In data and analytics — what you may do with what you can do, and how to keep the people in your rows safe.
Knowledge Discovery in Data · University of the Philippines Cebu
Real harms, famous failures, consent — and the law that binds you here.
Identifiers, k-anonymity, generalization — auditing and reducing re-identification risk.
Bias, transparency, and the tests a professional applies to their own work.
Week 4 gave you the scraping-etiquette preview and promised the full treatment later. This is later. You can now run the whole KDD pipeline — so the question "should you?" finally gets its own week.
Every row is a person who never met you. Act like they're watching.
Any data that identifies a person, on its own or when combined with other data.
e.g. a name, a student number, or age + barangay + course together
Personal information the law protects more strictly: health, education, government IDs, religion…
e.g. a student's GPA is education data, so it is sensitive
"Personally identifiable information": the common industry name for roughly the same idea as personal information.
e.g. email address, phone number, home address
The person the data is about. The law gives them rights over it.
e.g. the student in row 3 of a class list
A person's clear, informed "yes" to their data being used for a stated purpose.
e.g. ticking "use my answers to evaluate this course"
Anything done with personal data: collecting, storing, analysing, sharing, deleting.
e.g. running a groupby on a class survey is processing
Data collected for one purpose is used for that purpose only.
e.g. enrolment phone numbers are not a marketing list
Collect and keep only the columns the purpose really needs.
e.g. birth year instead of full birth date
A column that names one person by itself.
e.g. student_id, full name, email
A column that is harmless alone but can point to one person when combined with others.
e.g. age, barangay, birth date, sex
Matching an "anonymous" record back to a real, named person.
e.g. joining a health table to a voter list on ZIP + birth date + sex
Every combination of quasi-identifiers is shared by at least k records. k = 1 means someone stands alone.
e.g. k = 3: everyone hides in a crowd of at least three
Making a column less precise so more records look alike.
e.g. age 19 → "18–20"; exact budget → ₱50M band
Withholding any number or row that describes too few people.
e.g. no average published for a group of fewer than 5
Changing data so that nobody in it can be identified, by anyone, any more. Hard to achieve.
e.g. only group totals released, every group large
Replacing names with a code. Rows still link together, so it is not anonymous.
e.g. 2024-001 → 3e0f221a8d0c
A systematic tilt in data or analysis, so results lean one way for some group.
e.g. an online survey that misses people without internet
Whether an analysis or decision treats different groups of people equally well.
e.g. the same error rate for rich and poor barangays
Data about people carries risk — for them, and for you.
You can already collect, clean, explore and present data (weeks 3–10). This part asks the question before all of those: what the people in the rows are owed.
To your notebook, a row is four values. To its owner, it might be a location an abuser shouldn't learn, a diagnosis an employer shouldn't see, a grade a family shouldn't shame.
Stalking, discrimination in hiring or insurance, public humiliation, identity theft — all documented consequences of leaked datasets.
You publish once; the person carries it forever. Data spreads and never fully un-spreads.
The burden of care sits on the analyst — the one person who sees the whole table.
"Anonymized" queries for 650k users. Reporters identified user 4417749 by her own searches within days. Searches are a diary.
Movie ratings, no names. Researchers matched rating patterns to public IMDb reviews and re-identified users — patterns are fingerprints.
Aggregated jogging routes revealed the layouts of military bases. Even aggregates leak when the group is small and distinct.
Common thread: everyone involved believed the data was safe to release. "We removed the names" is where each failure began. Anonymised means nobody in the data can be identified any more; each release was only called that, and each was re-identified: records matched back to named people. An aggregate is a summary over a group, such as a total or an average.
Turn "we removed the names" into an actual audit you can run in pandas.
| Case | What was released | How it broke |
|---|---|---|
| Sweeney, 1997 | State employee health records with names removed | ZIP + birthdate + sex uniquely identified most people; matched against a public voter roll |
| AOL, 2006 | 20 million search queries, user IDs replaced with numbers | The searches themselves were identifying — journalists traced user 4417749 to a named individual |
| Netflix Prize, 2006 | 100 million movie ratings, subscriber IDs removed | Matched against public IMDb reviews; a handful of dated ratings identified subscribers |
The pattern is identical every time: the identifier was removed and the data was still identifying, because it could be joined against something public. (A ZIP code is a US postcode; a voter roll is the public list of registered voters, with names and birth dates.) Sweeney’s work is why k-anonymity exists — and why Part B computes it on our own dataset. The Netflix data came out in 2006; the re-identification paper followed in 2008.
Each release was safe against the data available that day. None survived a join with something that appeared later. (A quasi-identifier, in the code comment, is a column that is harmless alone but identifying in combination; Part B makes it precise.)
Data minimisation means collecting only what the purpose needs. A column you never collected cannot be joined against anything. A column you "anonymised" can be, by data that does not exist yet.
Consent is a person's clear, informed "yes" to a stated use of their data. People hand over data for something — enrollment, delivery, a health consult. Reusing it for something else is a new decision that needs its own justification.
Consent is like lending your car "to go to the market". It does not cover a road trip to Manila, however carefully the borrower drives.
Collected for X, used for X. The enrollment form's phone numbers are not a marketing list.
Collect the minimum that serves the purpose. Every extra column is extra risk you chose to hold.
Week 4's lesson, now sharpened: a reachable post is not permission to compile a profile. Aggregating public traces creates something new — and more dangerous.
Philippine law, enforced by the National Privacy Commission. If your rows describe people, this act describes your obligations.
| Concept | Meaning |
|---|---|
| Personal information | Data that identifies a person, alone or combined |
| Sensitive personal info | Health, education, government IDs, and more — stricter rules |
| Data subject rights | To be informed, to access, to correct, to object |
| Accountability | Whoever processes the data answers for its protection |
Education records — like GPAs — are sensitive personal information under the act. Today's lab dataset is regulated data in miniature.
The Philippine Data Privacy Act of 2012 is not a general principle — it imposes specific, checkable duties on anyone processing personal data.
Scope is about processing personal information, not about being a company. A class survey with names in it is covered.
Treat personal data like cash a friend handed you for one errand: use it for that errand, take only what you need, keep it safe, and give it back or destroy it when the errand is done. (This is a summary for analysts, not legal advice.)
Consent given for one stated purpose does not extend to the next idea you have. Re-using data for a new purpose needs a new basis.
Would the person have agreed to this use, at the moment they said yes? If you are not sure, you need to ask again.
For a study, you scrape thousands of public social-media posts with usernames and publish the dataset. Your defense: "It was all public." Does that hold?
C — collection creates something the posts alone were not.
A thousand scattered posts and one searchable dataset of named people are different objects with different risks. Politeness (B) answers the server-load question, not the persons-in-the-data question. At minimum: de-identify, aggregate, and know your legal basis.
Week 4 asked "may I take it?" This week's harder question: "what am I creating when I combine it?"
Justify the dataset you're building, not just the requests you're sending.
Then: privacy as an engineering problem you can actually test.
Audit the risk. Then shrink it. In pandas.
You already group rows and count them
(groupby, weeks 6–7). This part uses exactly that to measure how exposed each record is,
and then to make it safer.
The danger isn't only in the obvious column. It's in the combination of innocent-looking ones around it.
student_id — direct identifier, names a person.
age, barangay — quasi-identifiers, harmless
alone, identifying together. gpa — sensitive: the thing
worth protecting.
"Latanya Sweeney showed that 87% of Americans are uniquely identified by just ZIP code, birth date and sex — and famously re-identified the governor of Massachusetts in an 'anonymized' medical dataset using exactly those columns."Quasi-identifiers + one outside dataset = a name
The attack recipe never changes: take the "anonymous" table, join it to a
public one (voter rolls, social media, a class list) on the quasi-identifiers, and the
names come back. It's week 6's merge (joining two tables on shared columns) — used as a
weapon.
"A 21-year-old from Apas in this class" is not a name, but in a small class it points at exactly one person, just as surely as a name would.
Which is exactly why you're the right person to run the audit first.
Group by the quasi-identifiers and check sizes. A group of 1 is a person who can be singled out. k-anonymity: a table is "k-anonymous" when every combination appears at least k times.
df.groupby(["age", "barangay"])put rows with the same age and barangay in one group.size()count the rows in each group19 Lahug 3three students share "19, Lahug": each hides among three21 Apas 1a group of one: that student is exposed, so k = 1The 21-year-old from Apas is unique — anyone who knows those two facts now knows their GPA. Smallest group = 1, so this table isn't even 2-anonymous.
Watch what one extra column does.
Group by region and the smallest crowd is 180 records. No single project can be picked out; every row has at least 179 others that look identical. (These rows are projects, not people; the audit works the same way on either.)
The smallest group size is the k. Higher is safer, and 180 is very comfortable.
len(g)how many groups: 18 regionsg.min()the smallest group's size, which is the k(g == 1).sum()g == 1 gives True/False per group; .sum() counts the Trues (True counts as 1)| Quasi-identifiers | Groups | Records alone (k=1) | % unique |
|---|---|---|---|
| region | 18 | 0 | 0.0% |
| region + year | 180 | 1 | 0.0% |
| region + province + year | 1,944 | 96 | 0.3% |
| + budget rounded to ₱1M | 16,646 | 10,534 | 30.9% |
Nobody would call a rounded budget figure identifying. Adding it takes the dataset from nothing unique to nearly a third of records unique. This is why quasi-identifiers are dangerous: each is innocuous and the combination is not.
The same audit runs for every column list, so wrap it in a function (week 2: a named, reusable piece of work). Then each experiment is one line.
def audit(cols):define a function called audit that takes a list of column namesalone = (g == 1).sum()how many groups hold exactly one recordprint(f"{len(g):,} ...")an f-string: :, adds thousands commas, :.1% shows a share as a percentage(df.budget/1e6).round()budget in millions of pesos, rounded: ₱12,345,678 → 12.030.9%almost a third of all projects are alone in their groupGeneralisation means making a column less precise so more records look alike. Keep the budget column, but in ₱50M bands instead of exact millions: uniqueness falls from 30.9% to 2.7%. Generalise place as well (province up to region) and it falls to 0.4%.
You can no longer say what a project cost to the peso, or which province it was in. You can still say which band and region it sat in — which is enough for most questions.
Like a photo taken from further away: faces blur into a crowd, but you can still count the crowd and see where it stands.
Every generalisation that protects someone also removes something a legitimate analysis could have used. Pretending otherwise is how both get done badly.
If your analysis compares regional totals, ₱50M bands cost you nothing. If it studies cost variation within a province, they gut it — so you need a different protection, not a different rounding.
Remove student_id outright. Then blur the quasi-identifiers —
exact age becomes an age band — so the crowds grow.
df.drop(columns=["student_id"])a copy of the table without that columnpd.cut(safe["age"], bins=[17, 20, 25], ...)sort each age into a band; the bins are the edges: over 17 up to 20, over 20 up to 25labels=["18-20", "21-25"]the names the two bands get"19, Lahug" and "20, Lahug" merge into "18–20, Lahug" — one bigger crowd instead of two small ones. Larger k, lower risk, same broad analytical shape.
Bands lose detail — a study of "GPA by exact age" dies when age becomes two bins. That loss is the price of protection, and someone must decide how much to pay.
Risk survives; the k=1 cells are still there under prettier labels.
The dataset becomes decorative — safe and useless.
Not the analyst alone. Purpose, audience and the data subjects' stakes all weigh in — that's why review processes exist.
You'd never publish the raw table — but a summary table feels safe. Check the counts: the "average GPA in Apas" is one specific student's GPA, published.
df.groupby("barangay")["gpa"]group the students by barangay and look at the gpa column.agg(["mean", "count"])two summaries per group: the average and how many students\ at the end of a line"this line continues on the next one"1.25, 1an average over one student: it is that student's GPASuppression means withholding any number that describes too few people. Statistical agencies suppress any cell below a minimum count (often 3–10) for exactly this reason. Adopt it: no statistic for a group smaller than your threshold.
Cordillera Administrative Region, 2026: one project. Publish "the average project in CAR in 2026 cost X" and you have published that single contract’s exact value.
c.nsmallest(3)the three smallest counts; a blank region means “same as the row above”safe = c[c >= 5]keep only cells with 5 or more records: suppression(180, 179)180 cells before, 179 after: one cell was withheldA statistic over a group of one is the individual. Over a group of two, either party can subtract themselves and get the other.
Suppress any group below a threshold — commonly 5 or 10. Publish the count of suppressed cells so the table is honest.
If you publish the total and all groups but one, the missing one is arithmetic. Suppress a second cell to protect the first.
Two overlapping aggregations can be differenced to recover an individual. This is what differential privacy formalises.
The second rule catches people constantly: suppressing exactly one cell while publishing the total tells the reader precisely what you tried to hide.
Suppression stops one bad query. It does not stop someone asking two overlapping questions and subtracting. Differential privacy attacks that directly: it adds a little random noise to every answer.
Like a survey where each person secretly flips a coin before answering: no single answer can be trusted, so nobody is exposed, yet the overall percentage still comes out close.
The answer you get is almost exactly as likely whether or not any single record is in the dataset. So the answer cannot reveal that record.
laplace_noise(scale=1/eps)a stand-in name, not a real pandas function: "a random number, usually small, positive or negative"epsepsilon, the privacy budget: how much each answer may revealnoisythe number you publish, e.g. 1,261 or 1,265 instead of exactly 1,263Small cells in a published table. Simple, and the reader can see what was withheld.
Releasing a dataset others will analyse. Wide bands, k above your threshold. This is what you did with the budget column.
A query service answering many questions over time. The only one that survives repeated, overlapping queries.
They are not competing techniques; they answer different threats. Choose by asking how the data will be accessed, not by which sounds most rigorous.
Hashing turns any text into a fixed-length scramble of
characters (a digest); SHA-256 is one standard recipe. Replacing student_id
with its digest keeps rows linkable across files without showing the id. Useful — but that's
pseudonymisation (a code instead of a name), not anonymisation.
The b before the quotes turns the text into bytes, which hashing needs; [:12] keeps the first 12 of 64 characters.
A pseudonym is a nickname. It hides the name, but everything done under the nickname still links up, and anyone with the list of nicknames can undo it.
The same input always gives the same hash — so anyone who can guess the id space (four-digit years + three digits?) can hash every candidate and match.
The whole point of a stable pseudonym is that rows still connect. Connection is exactly what re-identification attacks exploit.
Pseudonymized data is still personal data under privacy law. The obligations don't hash away.
You removed student_id before
publishing the class dataset. A classmate says it's now anonymous. Is it?
B — run the audit before you claim the word.
"We removed the names" is where AOL and Netflix started, too. Group by the quasi-identifiers, find the k=1 cells, generalize until the crowds are big enough — then talk about anonymity, carefully. (And D is a counsel of despair, not an argument.)
Anonymity is a property you verify, not a label you declare.
The k-anonymity audit is three lines of pandas. Run it on anything you share.
Bias, transparency, and the tests you apply to yourself.
You can now measure and reduce the risk to the people in a table. This part asks the wider questions: who the data leaves out, and whether the analysis is fair to them.
Week 5's "meaningful gaps" at population scale: surveys miss the unconnected, complaint data misses those who've given up complaining.
Records of past decisions encode past discrimination. Analyze them naively and you launder yesterday's bias into today's "objective finding."
Drop the sensitive column and its correlates remain — barangay can proxy for income, school for class. Week 8's correlations, working against you.
Bias is a systematic tilt, so results lean one way for some group; fairness asks whether an analysis treats different groups equally well. A proxy is a column that stands in for another (barangay for income). None of this says "don't analyze." It says: know who your data speaks for, who it's silent about, and say so in your findings.
The week-9 caveat discipline — "this data covers X, not everyone" is a caveat too.
"The data is biased" is not a finding. How the bias got in determines whether you can correct it, and how.
They need different fixes: reweighting (counting under-represented rows more heavily), better collection, or abandoning the question. Conflating them produces confident, wrong corrections.
Weighing only the people who came to the gym is selection bias. A bathroom scale that always reads 2 kg light is measurement bias. Either way the average is wrong, for different reasons.
| Question | If the answer is uncomfortable |
|---|---|
| Should this exist? What is the purpose, and does this dataset serve it? | Do not collect it. Minimisation is the only protection that cannot fail. |
| Who is exposed? Compute k on your quasi-identifiers. | Generalise until k clears your threshold, and record what that cost. |
| Who bears the risk? Is it the same people who get the benefit? | If the subjects carry the risk and others take the benefit, that asymmetry needs justifying out loud. |
| What if I am wrong? Who is harmed by a false claim here? | Raise your evidence bar to match. Named parties need more than aggregates do. |
None of these is answered by a library. They are judgements, and the professional obligation is to make them explicitly and in writing — so that someone can disagree with you before publication rather than after.
Cheap to run, and they catch most trouble before it starts.
If your methods appeared on the front page — described accurately — would you defend them or explain them away?
Could you look the person in row 3 in the eye and describe what you did with their data?
If you're constructing a justification, you already have your answer. Ask a supervisor, an ethics board, or the NPC's guidance — before, not after.
There is no personal privacy question here — it is public money and corporate entities. That does not make publication risk-free.
A false claim about a named firm is defamation (a false public statement that harms a reputation) rather than a privacy breach, but the obligation to get it right is identical.
Every field you keep is a field that can leak, be subpoenaed (ordered handed over by a court), or be re-identified later by data that does not exist yet. A retention period is how long you keep data before deleting it.
"What analysis needs this column?" If there is no answer, do not collect it. Retro-fitting privacy is far harder than not collecting.
| Location | Why it survives a delete | What to do |
|---|---|---|
| Backups | Restore points predate the deletion — by design | Know your retention window; time deletions against it |
| Derived tables | Aggregates and exports built before the delete | Track lineage: which outputs used this input? |
| Version control | A committed file lives in history forever | Never commit data; a later delete does not remove it |
| Other people’s copies | Anything shared or downloaded | Log who received what, and when |
Version control (e.g. git) keeps every saved version of your files; one saved version is a commit. Lineage is the record of which outputs were built from which inputs. "We deleted it" is usually "we deleted one copy". Before promising deletion to anyone, enumerate the copies — and note that the backups protecting you from data loss are also the reason deletion takes time.
Compute group sizes for region; then region+year; then +province; then +budget rounded to the million. Count k=1 groups at each step.
Which single extra column raises uniqueness most? Why that one?
Widen the band until fewer than 1% of records are unique. What resolution did that cost you?
Two sentences: what you generalised, and which analyses are still valid on the released version.
Step 4 is what a data steward actually produces. The protection is only useful if the next analyst knows what it did to their questions.
Grouping by region gives a minimum group size of 180. You add province, year, and budget rounded to the nearest million. What happens?
You publish mean project cost per region-year. One cell (Cordillera, 2026) contains a single project. What have you published?
Four students, one sensitive column. You'll classify the identifiers, find the k=1 group, fix it by dropping and banding, and catch an aggregate that leaks. ~45 minutes.
Identifier triage, groupby().size() as a k-anonymity audit,
pd.cut age bands, and count-based suppression.
Hash an id with hashlib.sha256 and argue pseudonymous-vs-anonymous; hunt
the riskiest quasi-identifier combo.
Purpose-bound consent, minimal collection — and RA 10173 makes it law, not etiquette.
Audit quasi-identifiers with a groupby; k=1 means someone's exposed.
Remove direct ids, band the quasi columns, never publish a group of one — and know hashing is only a pseudonym.
Say who the data covers and who it's silent about. That sentence belongs in every report.
One sentence: protect the people in the rows as carefully as you polish the findings about them.
Next week the whole course folds into one arc — and your presentations close it.
The National Privacy Commission's plain-language primer on RA 10173 — your
obligations and your rights, in one read. privacy.gov.ph
The story of Latanya Sweeney re-identifying Governor Weld's medical record — the paper that made k-anonymity a field.
Both are linked on the course page beside this deck and the lab.
Having found a k=1 group with your own hands, Sweeney's attack reads like a recipe you already know.
The whole KDD arc in one view — and your final presentations bring it home.
DS 227 · Knowledge Discovery in Data