DS 227 · Week 11

Ethics & Privacy

In data and analytics — what you may do with what you can do, and how to keep the people in your rows safe.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

The week the pipeline grows a conscience

Words for Today · 1 of 2

Eight words about the people in your data

personal information (personal data)

Any data that identifies a person, on its own or when combined with other data.

e.g. a name, a student number, or age + barangay + course together

sensitive personal information

Personal information the law protects more strictly: health, education, government IDs, religion…

e.g. a student's GPA is education data, so it is sensitive

PII

"Personally identifiable information": the common industry name for roughly the same idea as personal information.

e.g. email address, phone number, home address

data subject

The person the data is about. The law gives them rights over it.

e.g. the student in row 3 of a class list

consent

A person's clear, informed "yes" to their data being used for a stated purpose.

e.g. ticking "use my answers to evaluate this course"

processing

Anything done with personal data: collecting, storing, analysing, sharing, deleting.

e.g. running a groupby on a class survey is processing

purpose limitation

Data collected for one purpose is used for that purpose only.

e.g. enrolment phone numbers are not a marketing list

data minimisation

Collect and keep only the columns the purpose really needs.

e.g. birth year instead of full birth date

Words for Today · 2 of 2

Ten words for protecting them

direct identifier

A column that names one person by itself.

e.g. student_id, full name, email

quasi-identifier

A column that is harmless alone but can point to one person when combined with others.

e.g. age, barangay, birth date, sex

re-identification

Matching an "anonymous" record back to a real, named person.

e.g. joining a health table to a voter list on ZIP + birth date + sex

k-anonymity

Every combination of quasi-identifiers is shared by at least k records. k = 1 means someone stands alone.

e.g. k = 3: everyone hides in a crowd of at least three

generalisation

Making a column less precise so more records look alike.

e.g. age 19 → "18–20"; exact budget → ₱50M band

suppression

Withholding any number or row that describes too few people.

e.g. no average published for a group of fewer than 5

anonymisation

Changing data so that nobody in it can be identified, by anyone, any more. Hard to achieve.

e.g. only group totals released, every group large

pseudonymisation

Replacing names with a code. Rows still link together, so it is not anonymous.

e.g. 2024-001 → 3e0f221a8d0c

bias

A systematic tilt in data or analysis, so results lean one way for some group.

e.g. an online survey that misses people without internet

fairness

Whether an analysis or decision treats different groups of people equally well.

e.g. the same error rate for rich and poor barangays

Part A

Why It Matters

Data about people carries risk — for them, and for you.

You can already collect, clean, explore and present data (weeks 3–10). This part asks the question before all of those: what the people in the rows are owed.

The Stakes

A leaked row is a changed life

To your notebook, a row is four values. To its owner, it might be a location an abuser shouldn't learn, a diagnosis an employer shouldn't see, a grade a family shouldn't shame.

Harms are concrete

Stalking, discrimination in hiring or insurance, public humiliation, identity theft — all documented consequences of leaked datasets.

And asymmetric

You publish once; the person carries it forever. Data spreads and never fully un-spreads.

Which is why

The burden of care sits on the analyst — the one person who sees the whole table.

Case Files

Three releases that taught the field

Three Releases That Built This Field

All three were "anonymised" first

CaseWhat was releasedHow it broke
Sweeney, 1997State employee health records with names removedZIP + birthdate + sex uniquely identified most people; matched against a public voter roll
AOL, 200620 million search queries, user IDs replaced with numbersThe searches themselves were identifying — journalists traced user 4417749 to a named individual
Netflix Prize, 2006100 million movie ratings, subscriber IDs removedMatched against public IMDb reviews; a handful of dated ratings identified subscribers
The Lesson All Three Share

You cannot anonymise against future datasets

Each release was safe against the data available that day. None survived a join with something that appeared later. (A quasi-identifier, in the code comment, is a column that is harmless alone but identifying in combination; Part B makes it precise.)

Which is why minimisation beats anonymisation

Data minimisation means collecting only what the purpose needs. A column you never collected cannot be joined against anything. A column you "anonymised" can be, by data that does not exist yet.

# what each team believed "we removed the identifiers" # what was true "we removed the identifiers we thought of, and left a fingerprint that matched a public dataset we had not considered" # your quasi-identifier analysis is # only as good as the outside data # you imagined
First Principles

Consent is for a purpose, not forever

Consent is a person's clear, informed "yes" to a stated use of their data. People hand over data for something — enrollment, delivery, a health consult. Reusing it for something else is a new decision that needs its own justification.

Consent is like lending your car "to go to the market". It does not cover a road trip to Manila, however carefully the borrower drives.

Purpose limitation

Collected for X, used for X. The enrollment form's phone numbers are not a marketing list.

Proportionality

Collect the minimum that serves the purpose. Every extra column is extra risk you chose to hold.

Public ≠ consented

Week 4's lesson, now sharpened: a reachable post is not permission to compile a profile. Aggregating public traces creates something new — and more dangerous.

Not Optional

RA 10173: the Data Privacy Act of 2012

Philippine law, enforced by the National Privacy Commission. If your rows describe people, this act describes your obligations.

Words the act uses

data subject
the person the data is about
processing
anything done with personal data: collect, store, analyse, share, delete
NPC
National Privacy Commission, the government office that enforces the act
PII
"personally identifiable information", the industry name for personal information
ConceptMeaning
Personal informationData that identifies a person, alone or combined
Sensitive personal infoHealth, education, government IDs, and more — stricter rules
Data subject rightsTo be informed, to access, to correct, to object
AccountabilityWhoever processes the data answers for its protection

Note the second row

Education records — like GPAs — are sensitive personal information under the act. Today's lab dataset is regulated data in miniature.

RA 10173, Concretely

What the Data Privacy Act actually requires of you

The Philippine Data Privacy Act of 2012 is not a general principle — it imposes specific, checkable duties on anyone processing personal data.

It applies to student projects

Scope is about processing personal information, not about being a company. A class survey with names in it is covered.

Treat personal data like cash a friend handed you for one errand: use it for that errand, take only what you need, keep it safe, and give it back or destroy it when the errand is done. (This is a summary for analysts, not legal advice.)

• collect for a declared, specific purpose — and only that purpose • collect the minimum needed • keep it accurate and current • retain only as long as the purpose requires, then dispose securely • secure it — organisationally and technically • honour the subject's rights: to be informed, to access, to object, to erasure, to damages
Consent Is For A Purpose

Not a permanent licence

Consent given for one stated purpose does not extend to the next idea you have. Re-using data for a new purpose needs a new basis.

The test

Would the person have agreed to this use, at the moment they said yes? If you are not sure, you need to ask again.

# what they consented to "to evaluate this course" # what that does NOT cover - publishing quotes with names - sharing with another department - training a model (a prediction program learned from the data) - keeping it after the term ends # each is a new purpose needing # a new basis — usually fresh consent
Quick Check

Tap to reveal

For a study, you scrape thousands of public social-media posts with usernames and publish the dataset. Your defense: "It was all public." Does that hold?

A · Yes — public means fair game
B · Yes, if you scraped slowly and politely
C · No — compiling identified profiles is new processing of personal data, needing its own justification
D · Only lawyers can have an opinion

C — collection creates something the posts alone were not.

A thousand scattered posts and one searchable dataset of named people are different objects with different risks. Politeness (B) answers the server-load question, not the persons-in-the-data question. At minimum: de-identify, aggregate, and know your legal basis.

Break

  Five minutes

Then: privacy as an engineering problem you can actually test.

Part B

Privacy Engineering

Audit the risk. Then shrink it. In pandas.

You already group rows and count them (groupby, weeks 6–7). This part uses exactly that to measure how exposed each record is, and then to make it safer.

Know Your Columns

Direct, quasi, sensitive — three kinds of column

The danger isn't only in the obvious column. It's in the combination of innocent-looking ones around it.

# the lab's four students print(df) student_id age barangay gpa 0 2024-001 19 Lahug 1.75 1 2024-002 19 Lahug 2.00 2 2024-003 21 Apas 1.25 3 2024-004 19 Lahug 1.50

Classify them

student_id — direct identifier, names a person. age, barangay — quasi-identifiers, harmless alone, identifying together. gpa — sensitive: the thing worth protecting.

The Core Threat

Nobody needs your name to find you

"Latanya Sweeney showed that 87% of Americans are uniquely identified by just ZIP code, birth date and sex — and famously re-identified the governor of Massachusetts in an 'anonymized' medical dataset using exactly those columns."
Quasi-identifiers + one outside dataset = a name
The Audit

k-anonymity: count your crowds

Group by the quasi-identifiers and check sizes. A group of 1 is a person who can be singled out. k-anonymity: a table is "k-anonymous" when every combination appears at least k times.

  • df.groupby(["age", "barangay"])put rows with the same age and barangay in one group
  • .size()count the rows in each group
  • 19 Lahug 3three students share "19, Lahug": each hides among three
  • 21 Apas 1a group of one: that student is exposed, so k = 1
sizes = df.groupby( ["age", "barangay"]).size() print(sizes) age barangay 19 Lahug 3 21 Apas 1 ← k=1: exposed dtype: int64

Read the audit

The 21-year-old from Apas is unique — anyone who knows those two facts now knows their GPA. Smallest group = 1, so this table isn't even 2-anonymous.

Compute It, Do Not Just Define It

k-anonymity on 34,079 real records

Watch what one extra column does.

Start Safe

Region alone hides everyone

Group by region and the smallest crowd is 180 records. No single project can be picked out; every row has at least 179 others that look identical. (These rows are projects, not people; the audit works the same way on either.)

k = 180

The smallest group size is the k. Higher is safer, and 180 is very comfortable.

  • len(g)how many groups: 18 regions
  • g.min()the smallest group's size, which is the k
  • (g == 1).sum()g == 1 gives True/False per group; .sum() counts the Trues (True counts as 1)
g = df.groupby(["region"]).size() len(g) 18 # groups g.min() 180 # k = 180 (g == 1).sum() 0 # nobody is alone
Now Add Columns, One At A Time

Each one looks harmless on its own

Quasi-identifiersGroupsRecords alone (k=1)% unique
region1800.0%
region + year18010.0%
region + province + year1,944960.3%
+ budget rounded to ₱1M16,64610,53430.9%
An Audit You Can Reuse

Three lines, wrapped in a function

The same audit runs for every column list, so wrap it in a function (week 2: a named, reusable piece of work). Then each experiment is one line.

  • def audit(cols):define a function called audit that takes a list of column names
  • alone = (g == 1).sum()how many groups hold exactly one record
  • print(f"{len(g):,} ...")an f-string: :, adds thousands commas, :.1% shows a share as a percentage
  • (df.budget/1e6).round()budget in millions of pesos, rounded: ₱12,345,678 → 12.0
  • 30.9%almost a third of all projects are alone in their group
def audit(cols): g = df.groupby(cols).size() alone = (g == 1).sum() print(f"{len(g):,} groups, {alone:,} alone " f"({alone/len(df):.1%})") qi = ["region", "province", "year", "b"] # precise: PHP millions df["b"] = (df.budget/1e6).round() audit(qi) 16,646 groups, 10,534 alone (30.9%)
The Same Move, Undone

Generalisation buys the privacy back

Generalisation means making a column less precise so more records look alike. Keep the budget column, but in ₱50M bands instead of exact millions: uniqueness falls from 30.9% to 2.7%. Generalise place as well (province up to region) and it falls to 0.4%.

You lost precision to gain it

You can no longer say what a project cost to the peso, or which province it was in. You can still say which band and region it sat in — which is enough for most questions.

Like a photo taken from further away: faces blur into a crowd, but you can still count the crowd and see where it stands.

# generalised: PHP 50M bands df["b"] = (df.budget/5e7).round() * 50 audit(qi) 4,675 groups, 915 alone (2.7%) # and generalise place: region, not province audit(["region", "year", "b"]) 821 groups, 132 alone (0.4%) # bands alone: 11x fewer singletons. # bands + region: 80x fewer.
The Trade-Off Is Real

You cannot have full precision and full privacy

Every generalisation that protects someone also removes something a legitimate analysis could have used. Pretending otherwise is how both get done badly.

Decide by the question

If your analysis compares regional totals, ₱50M bands cost you nothing. If it studies cost variation within a province, they gut it — so you need a different protection, not a different rounding.

# the honest framing precision ────────────► privacy exact values wide bands, by province by region 30.9% unique 0.4% unique any analysis aggregate only pick the LEAST generalisation that gets k above your threshold — not the most you can get away with
The Treatment

Drop the direct, generalize the quasi

Remove student_id outright. Then blur the quasi-identifiers — exact age becomes an age band — so the crowds grow.

  • df.drop(columns=["student_id"])a copy of the table without that column
  • pd.cut(safe["age"], bins=[17, 20, 25], ...)sort each age into a band; the bins are the edges: over 17 up to 20, over 20 up to 25
  • labels=["18-20", "21-25"]the names the two bands get
  • Re-run the audit(18–20, Lahug) has 3; (21–25, Apas) still has 1. Banding age alone did not save the Apas student: next comes suppression
safe = df.drop(columns=["student_id"]) safe["age_band"] = pd.cut(safe["age"], bins=[17, 20, 25], labels=["18-20", "21-25"]) safe = safe.drop(columns=["age"])

What generalization buys

"19, Lahug" and "20, Lahug" merge into "18–20, Lahug" — one bigger crowd instead of two small ones. Larger k, lower risk, same broad analytical shape.

No Free Lunch

Privacy and precision pull against each other

Bands lose detail — a study of "GPA by exact age" dies when age becomes two bins. That loss is the price of protection, and someone must decide how much to pay.

Blur too little

Risk survives; the k=1 cells are still there under prettier labels.

Blur too much

The dataset becomes decorative — safe and useless.

Who draws the line?

Not the analyst alone. Purpose, audience and the data subjects' stakes all weigh in — that's why review processes exist.

The Subtle Leak

A mean of one is not a mean

You'd never publish the raw table — but a summary table feels safe. Check the counts: the "average GPA in Apas" is one specific student's GPA, published.

  • df.groupby("barangay")["gpa"]group the students by barangay and look at the gpa column
  • .agg(["mean", "count"])two summaries per group: the average and how many students
  • \ at the end of a line"this line continues on the next one"
  • 1.25, 1an average over one student: it is that student's GPA
df.groupby("barangay")["gpa"]\ .agg(["mean", "count"]) mean count barangay Apas 1.25 1 ← leak Lahug 1.75 3

The suppression rule

Suppression means withholding any number that describes too few people. Statistical agencies suppress any cell below a minimum count (often 3–10) for exactly this reason. Adopt it: no statistic for a group smaller than your threshold.

A Mean Of One

The smallest real cell in this dataset

Cordillera Administrative Region, 2026: one project. Publish "the average project in CAR in 2026 cost X" and you have published that single contract’s exact value.

  • c.nsmallest(3)the three smallest counts; a blank region means “same as the row above”
  • safe = c[c >= 5]keep only cells with 5 or more records: suppression
  • (180, 179)180 cells before, 179 after: one cell was withheld

Aggregation is not automatically anonymisation

A statistic over a group of one is the individual. Over a group of two, either party can subtract themselves and get the other.

c = df.groupby(["region", "year"]).size() c.nsmallest(3) region year Cordillera Administrative Region 2026 1 Central Office 2024 6 2020 13 dtype: int64 # suppress cells below a threshold safe = c[c >= 5] len(c), len(safe) (180, 179)
Three Rules For Publishing Aggregates

What statistical agencies actually do

Differential Privacy

The answer to the repeated-query problem

Suppression stops one bad query. It does not stop someone asking two overlapping questions and subtracting. Differential privacy attacks that directly: it adds a little random noise to every answer.

Like a survey where each person secretly flips a coin before answering: no single answer can be trusted, so nobody is exposed, yet the overall percentage still comes out close.

The guarantee, in words

The answer you get is almost exactly as likely whether or not any single record is in the dataset. So the answer cannot reveal that record.

# the mechanism: add calibrated noise true_count = 1263 noisy = true_count + laplace_noise(scale=1/eps) # small eps -> more noise, more privacy # large eps -> less noise, less privacy # and the budget is CUMULATIVE: # every query you answer spends # some of it. that is the point — # it makes the difference attack # cost something
  • laplace_noise(scale=1/eps)a stand-in name, not a real pandas function: "a random number, usually small, positive or negative"
  • epsepsilon, the privacy budget: how much each answer may reveal
  • noisythe number you publish, e.g. 1,261 or 1,265 instead of exactly 1,263
When To Reach For What

Three protections, three situations

A Common Half-Measure

Hashing hides the name, not the person

Hashing turns any text into a fixed-length scramble of characters (a digest); SHA-256 is one standard recipe. Replacing student_id with its digest keeps rows linkable across files without showing the id. Useful — but that's pseudonymisation (a code instead of a name), not anonymisation.

import hashlib hashlib.sha256(b"2024-001").hexdigest()[:12] '3e0f221a8d0c'

The b before the quotes turns the text into bytes, which hashing needs; [:12] keeps the first 12 of 64 characters.

A pseudonym is a nickname. It hides the name, but everything done under the nickname still links up, and anyone with the list of nicknames can undo it.

Why it's not anonymity

The same input always gives the same hash — so anyone who can guess the id space (four-digit years + three digits?) can hash every candidate and match.

And linkability persists

The whole point of a stable pseudonym is that rows still connect. Connection is exactly what re-identification attacks exploit.

Legal echo

Pseudonymized data is still personal data under privacy law. The obligations don't hash away.

Quick Check

Tap to reveal

You removed student_id before publishing the class dataset. A classmate says it's now anonymous. Is it?

A · Yes — no id, no identification
B · Not necessarily — quasi-identifier combos with k=1 can still single people out
C · Yes, if the file is renamed "anonymous.csv"
D · Anonymity is impossible, so publish anything

B — run the audit before you claim the word.

"We removed the names" is where AOL and Netflix started, too. Group by the quasi-identifiers, find the k=1 cells, generalize until the crowds are big enough — then talk about anonymity, carefully. (And D is a counsel of despair, not an argument.)

Part C

Beyond Privacy

Bias, transparency, and the tests you apply to yourself.

You can now measure and reduce the risk to the people in a table. This part asks the wider questions: who the data leaves out, and whether the analysis is fair to them.

The Other Ethical Front

Data inherits the world's unfairness

Bias Has A Mechanism

Name it, or you cannot fix it

"The data is biased" is not a finding. How the bias got in determines whether you can correct it, and how.

Three distinct mechanisms

They need different fixes: reweighting (counting under-represented rows more heavily), better collection, or abandoning the question. Conflating them produces confident, wrong corrections.

Weighing only the people who came to the gym is selection bias. A bathroom scale that always reads 2 kg light is measurement bias. Either way the average is wrong, for different reasons.

# 1. SELECTION — who got in only completed projects have a duration -> your "typical project" excludes everything that stalled # 2. MEASUREMENT — how it was recorded amount_paid is never populated -> any "unpaid" analysis is an artifact of the recording process # 3. HISTORICAL — the world was unfair past allocation patterns encode past priorities; a model fit on them reproduces those priorities
The Decision Framework

Four questions, before you collect, keep or publish

QuestionIf the answer is uncomfortable
Should this exist? What is the purpose, and does this dataset serve it?Do not collect it. Minimisation is the only protection that cannot fail.
Who is exposed? Compute k on your quasi-identifiers.Generalise until k clears your threshold, and record what that cost.
Who bears the risk? Is it the same people who get the benefit?If the subjects carry the risk and others take the benefit, that asymmetry needs justifying out loud.
What if I am wrong? Who is harmed by a false claim here?Raise your evidence bar to match. Named parties need more than aggregates do.
Carrying It Forward

Three tests before you collect, keep or publish

Cheap to run, and they catch most trouble before it starts.

  The front-page test

If your methods appeared on the front page — described accurately — would you defend them or explain them away?

  The row-3 test

Could you look the person in row 3 in the eye and describe what you did with their data?

  The hesitation test

If you're constructing a justification, you already have your answer. Ask a supervisor, an ethics board, or the NPC's guidance — before, not after.

Public Data Is Not Consequence-Free

This dataset names real companies

There is no personal privacy question here — it is public money and corporate entities. That does not make publication risk-free.

Different harm, same care

A false claim about a named firm is defamation (a false public statement that harms a reputation) rather than a privacy breach, but the obligation to get it right is identical.

# the week 9 example, as an ethics case 131 contractors carry "[REVOKED]" 1,263 projects, PHP 75.2B The data supports: "flagged licence-revoked in the DPWH export" The data does NOT support: "awarded contracts while their licence was revoked" — there is no revocation date.
Minimise What You Collect

The safest record is the one you never stored

Every field you keep is a field that can leak, be subpoenaed (ordered handed over by a court), or be re-identified later by data that does not exist yet. A retention period is how long you keep data before deleting it.

Ask the question at collection time

"What analysis needs this column?" If there is no answer, do not collect it. Retro-fitting privacy is far harder than not collecting.

# do you need... exact birthdate? -> birth year full address? -> barangay exact coordinates? -> municipality name + email? -> a random id free-text comments? -> often not # and set a retention period at the # same moment you create the field
Deletion Is Harder Than It Looks

Four places the data still is

LocationWhy it survives a deleteWhat to do
BackupsRestore points predate the deletion — by designKnow your retention window; time deletions against it
Derived tablesAggregates and exports built before the deleteTrack lineage: which outputs used this input?
Version controlA committed file lives in history foreverNever commit data; a later delete does not remove it
Other people’s copiesAnything shared or downloadedLog who received what, and when
Your Turn · 8 min

Re-identify, then protect

1 · Reproduce the progression

Compute group sizes for region; then region+year; then +province; then +budget rounded to the million. Count k=1 groups at each step.

2 · Find your own quasi-identifier

Which single extra column raises uniqueness most? Why that one?

3 · Generalise it back

Widen the band until fewer than 1% of records are unique. What resolution did that cost you?

4 · Write the release note

Two sentences: what you generalised, and which analyses are still valid on the released version.

Quick Check

Tap to reveal

Grouping by region gives a minimum group size of 180. You add province, year, and budget rounded to the nearest million. What happens?

A · Little changes — rounded budgets are not identifying
B · 30.9% of records become uniquely identifiable
C · Privacy improves, because the groups are more specific
D · It has no effect unless you also include names
B. Measured: 16,646 groups, 10,534 of them containing a single record — from 0 singletons with region alone. No individual column here is identifying; the combination is. That is precisely what a quasi-identifier is, and why privacy review has to consider column sets rather than columns.
Quick Check

Tap to reveal

You publish mean project cost per region-year. One cell (Cordillera, 2026) contains a single project. What have you published?

A · An aggregate, which is inherently anonymous
B · That project’s exact cost — a mean over one record is the record
C · Nothing sensitive, since it is public spending data
D · A biased estimate, but nothing identifying
B. Aggregation protects only when the group is large enough to hide in. A mean over n=1 is the value itself; over n=2, either party can subtract their own and recover the other. This is why agencies suppress small cells — and why you must also suppress a second cell, or the total gives the first one back by subtraction.
This Week's Lab

Audit a dataset for privacy risk

Four students, one sensitive column. You'll classify the identifiers, find the k=1 group, fix it by dropping and banding, and catch an aggregate that leaks. ~45 minutes.

You'll practise

Identifier triage, groupby().size() as a k-anonymity audit, pd.cut age bands, and count-based suppression.

Stretch, if you're quick

Hash an id with hashlib.sha256 and argue pseudonymous-vs-anonymous; hunt the riskiest quasi-identifier combo.

Recap

Four things to carry out

Glossary

This week's words, one line each

Words for today

personal information (personal data)
data that identifies a person, alone or combined
sensitive personal information
health, education, IDs…: stricter rules
PII
personally identifiable information (industry term)
data subject
the person the data is about
consent
an informed "yes" to one stated use
processing
anything done with personal data
purpose limitation
collected for X, used for X only
data minimisation
collect and keep only what the purpose needs
direct identifier
a column that names a person by itself
quasi-identifier
harmless alone, identifying in combination
re-identification
matching an "anonymous" row to a named person
k-anonymity
every combination shared by at least k records
generalisation
less precise values so more rows look alike
suppression
withholding numbers about too few people
anonymisation
nobody can be identified, by anyone, any more
pseudonymisation
a code instead of a name; rows still link
bias
a systematic tilt for or against some group
fairness
treating different groups equally well

Also new today

RA 10173
the Philippine Data Privacy Act of 2012
NPC
National Privacy Commission, which enforces it
data subject rights
to be informed, access, correct, object…
sensitive (column)
the value worth protecting, e.g. GPA
aggregate
a summary over a group: total, average
differential privacy
small random noise added to every answer
hashing / digest
a fixed scramble of text; same input, same output
proxy
a column that stands in for another
selection / measurement bias
who got in / how it was recorded
retention period
how long data is kept before deletion
defamation
a false public claim that harms a reputation
lineage
which outputs were built from which inputs
Readings

Before next week

Next Week

Course Synthesis

The whole KDD arc in one view — and your final presentations bring it home.

DS 227 · Knowledge Discovery in Data