DS 227 · Week 9

Data Journalism

Finding stories in data — turning a question into a comparison, a comparison into a chart, and a chart into one sentence you can defend.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

From dataset to defensible claim

Words for Today

Ten words a data journalist uses

story

A finding that a reader should know, told so that they can check it.

e.g. "48 of 4,841 firms hold 27.6% of the money"

angle

The particular question or point of view a story takes on the data. One dataset has many angles.

e.g. flood data: spending per region, or which firms get the most

comparison

Setting a number next to another group, time or expectation. Most findings are one.

e.g. passers studied 7.0 h; the others 2.4 h

context

The surrounding facts a reader needs to judge whether a number is big, small or normal.

e.g. Region IV-B builds different works than Region I

denominator

The number you divide by: the "out of how many" or "per what" behind a figure.

e.g. ₱1.6T ÷ 11 years = ₱145.5B per year

per-capita

Per person: divide a count by the population, so places of different sizes compare fairly.

e.g. 500 cases ÷ 1,000,000 people = 50 per 100,000

base rate

How common something is overall: the yardstick you compare a group against.

e.g. 0.4% of all projects are terminated

source

Where the data came from: who produced it and where you got it.

e.g. DPWH flood control projects, via the BetterGov.ph mirror

verification

Checking a claim against the evidence before publishing it.

e.g. look at .unique() on the column a claim rests on

unpopulated column

A column that exists but was never filled in. Its zeros mean "not recorded", not "none". The amount_paid trap.

e.g. df.amount_paid.unique() gives array([0])

Part A

Stories Live In Data

They always have — since before "data science" was a phrase.

For six weeks you have collected, cleaned and explored data. This part asks what turns a pattern you found into a story someone else should read.

The Original Data Story

1854: a map finds a pump

"During London's cholera outbreak, John Snow plotted each death as a mark on a street map. The marks piled up around one water pump on Broad Street. The dominant theory said disease traveled by air; his picture said it traveled by water — and the pump handle came off."
One question, one dataset, one chart — and a changed world
Story Shapes

Three shapes a data story takes

One Dataset, Many Stories

Pick an angle before you pick a chart

The Starting Point

A story starts with a question, not a dataset

"Here's a CSV, find something" produces noise. "Does study time relate to passing?" names an outcome and a suspected driver — now every step has a direction.

The outcome is what you want to explain (passing); the driver is what you suspect explains it (study hours). A boolean column holds only True or False.

A good question names

The outcome you care about (passing) and the candidate explanation (study hours).

Then shape the data to it

The lab turns score into a boolean passed — the question got easier to ask, at the cost of nuance (a 74 and a 40 become the same "failed").

Name your trade-offs

Every simplification you make is one a critic can point at. Better you point at it first.

You Have Already Nearly Published Five Falsehoods

All on the same dataset

Weeks 4 through 8. Each one was a single keystroke away.

The Running Tally

Every week, a plausible claim that was wrong

WeekThe tempting claimWhy it was false
4₱1.17T of work completed but never paidamount_paid is 0 on every row — unpopulated, not unpaid
599.6% of flood projects are completedropna() deleted every unfinished project; the real figure is 80.8%
6One project finished 17 days before it startedTrue — and it is a data error, not a scandal
7Projects take a median of 202 daysdescribe() silently excluded 6,415 unfinished ones
8Latitude strongly predicts longituder = −0.785, and it is the shape of the archipelago
The amount_paid Trap

A column of zeros is not a column of "unpaid"

The flood data has an amount_paid column. Sum the budgets of completed projects, see that amount_paid is 0, and you have a ₱1.17 trillion "completed but never paid" scandal. Except the column is unpopulated: it exists, but nobody ever filled it in.

A blank "tip" line on a receipt does not prove nobody tipped; the restaurant may just not record tips on receipts.

The ten-second check

Before any claim rests on a column, look at its .unique() values. One value on every row means "not recorded" until proven otherwise.

df.amount_paid.unique() array([0]) df.amount_paid.sum() 0 # the tempting "scandal" number round(df[df.status == "Completed"].budget.sum() / 1e12, 2) 1.17 # PHP trillion
  • .unique()every different value in the column: here only 0
  • array([0])a NumPy array holding one value: every one of 34,079 rows says 0
  • df[df.status == "Completed"]keep the completed projects, then add up their budgets in trillions
The Pattern

The dangerous claim is the one that sounds like a story

Notice which of those five you would have been most excited to publish. The ₱1.17 trillion one. It was also the most wrong.

Excitement is a warning sign

A finding that immediately feels like a headline deserves more checking than one that feels mundane — not less. The story pulls you past the verification.

# the check that would have saved each one df.amount_paid.unique() # week 4 len(df), len(clean) # week 5 df.duration.min() # week 6 df.duration.describe()["count"] # week 7 why_would_these_relate() # week 8 # none took more than ten seconds
  • df.amount_paid.unique()list every different value in the column: one value only means nothing was recorded
  • len(df), len(clean)row count before and after cleaning: what did you drop?
  • why_would_these_relate()not real code: a question to ask yourself
Break

  Five minutes

Then: the craft — building the finding, step by step.

Part B

The Craft

Question → comparison → chart → sentence.

You know how to group and summarise (weeks 5–8). This part turns a group comparison into a chart and one sentence a reader can check.

Step 1

Turn the question into a comparison

Split the rows by the outcome, then summarize the candidate driver in each group. One groupby, and the question has numbers.

The data: 10 students with their weekly study hours and a passed column (True if the score was 75 or more).

df.groupby("passed")["study_hours"].agg( ["mean", "min", "max", "count"]) mean min max count passed False 2.4 1 4 5 True 7.0 5 9 5
  • groupby("passed")["study_hours"]split students into passed / not passed, and look at their study hours
  • .agg(["mean", "min", "max", "count"])four summaries at once for each group
  • True 7.0 5 9 5the 5 passers averaged 7 hours; the least any of them studied was 5

Read the whole row

Not just the means — the ranges don't even overlap here, and count tells you how much weight the claim can bear (not much, at 5 per group).

₱1.6 Trillion

Is that a lot?

You cannot answer that, and neither can your reader, without a denominator.

A Number Alone Means Nothing

Give it something to stand next to

₱1.6 trillion is the headline figure for this dataset. On its own it is unreadable — big enough to sound like everything, small enough to sound like nothing. A denominator is what you divide by to make it comparable: per year, per project, per person.

Divide by time first

It covers eleven years. ₱145.5B per year is a number someone can hold in their head and compare to other budget lines.

round(df.budget.sum() / 1e12, 2) 1.6 # PHP trillion, 2016-2026 df.year.nunique() 11 round(df.budget.sum() / 1e9 / 11, 1) 145.5 # PHP billion per year # still large. but now comparable.
  • df.budget.sum() / 1e12add every budget; 1e12 is a trillion, so the answer is in trillions (round keeps 2 decimals: 1.60 prints as 1.6)
  • df.year.nunique()how many different years appear: 11 (2016 to 2026)
  • … / 1e9 / 11in billions (1e9), then divided by 11 years
Three Denominators, Three Stories

Same ₱1.6T, different meanings

Per-Capita

Bigger places have more of everything

Per-capita means "per person". Divide a count by the population and places of very different sizes become comparable. The numbers on the right are made up to show the arithmetic.

The count and the rate disagree

City A has 12 times as many reports, yet Town B has the higher rate: 80 against 50 per 100,000 people. A headline built on the count would point at the wrong place.

A family of eight eating eight kilos of rice a week is not hungrier than a couple eating three. Divide by the people at the table.

# made-up example: flood reports and population reports = pd.Series({"City A": 500, "Town B": 40}) people = pd.Series({"City A": 1_000_000, "Town B": 50_000}) reports / people * 100_000 City A 50.0 Town B 80.0 dtype: float64
  • pd.Series({"City A": 500, …})a Series built from a dictionary: the names become the row labels
  • reports / peoplepandas lines the two Series up by label and divides City A by City A, Town B by Town B
  • * 100_000reports per 100,000 people, an easier number to read than 0.0008
  • dtype: float64the results are decimal numbers
Compared To What

The question a reader asks before believing you

Region IV-B spends ₱67.1M per project; Region I spends ₱31.2M. Twice as much — which sounds like a finding until you ask what kind of projects each builds. That is context: the surrounding facts a reader needs to judge whether a number is big, small or normal.

A difference is not yet a story

Bigger rivers, denser cities, and different works all produce different unit costs legitimately. The comparison generates the question; it does not answer it.

g = df.groupby("region").agg( n=("contract_id", "count"), total=("budget", "sum")) g["per_project"] = g.total / g.n top = (g.per_project / 1e6).round(1).sort_values(ascending=False) top.iloc[[0, 1, 2, -1]] # PHP millions region Central Office 242.0 Region IV-B 67.1 Region XIII 61.0 Region I 31.2 Name: per_project, dtype: float64
  • n=("contract_id", "count")a new column n: how many projects each region has
  • total=("budget", "sum")a new column total: the region's summed budget
  • g.total / g.ndivide one column by the other, row by row: average budget per project
  • top.iloc[[0, 1, 2, -1]]in millions, biggest first; show the top three and the last one (−1)
Step 2

Show the contrast — honestly

A bar chart of the two means makes the gap land in half a second. Note the y-axis: bars encode length, so they must start at zero — a truncated bar chart is a lie about proportion.

means = df.groupby("passed")[ "study_hours"].mean() ax.bar(["failed", "passed"], means) ax.set_ylabel("Avg study hours")
  • ax.bar(["failed", "passed"], means)two bars: the labels along the bottom, the two means as heights
  • ax.set_ylabel(…)say what the height measures; matplotlib starts bars at 0 unless you change it

Chart follows claim

The chart exists to carry this comparison — not to decorate. If it needs a paragraph to explain, the comparison isn't sharp yet.

What That Code Draws

Two bars from zero: 2.4 hours against 7.0

the y axis starts at 0, so the passed bar is really almost three times as tall as the failed bar (7.0 against 2.4), and your eye reads the true proportion. This is the lab's ten students, with fig, ax = plt.subplots() run first to make ax.

Bar chart with two bars labelled failed and passed: failed reaches 2.4 average study hours and passed reaches 7.0, on a y axis labelled Avg study hours that starts at 0.
Step 3

The finding is one sentence

Claim + number + caveat. A caveat is the honest limit you attach ("only 10 students"). If you can't say it in a sentence, you don't have a finding yet — you have an area of interest.

# the numbers from Step 1 # passed avg 7.0 h vs failed avg 2.4 h # gap 4.6 h, n = 10

The sentence

"Students who passed studied about 4.6 hours more per week on average — though with only 10 students, this is suggestive, not conclusive."

Anatomy

Claim (studied more) · number (4.6h, n=10) · caveat (suggestive). All three, every time.

On Hedging

The caveat is not weakness — it's accuracy

"An overclaimed finding gets one news cycle and then gets corrected in public. A precisely-claimed finding gets cited. 'Suggestive, not conclusive' isn't timidity — it's the exact size of the evidence."
Say what the data supports. Then stop talking.
Quick Check

Tap to reveal

Which is the defensible version of the lab's finding?

A · "Studying makes students pass."
B · "Study time matters a lot."
C · "Passers studied ~4.6h more on average (n=10) — suggestive, not conclusive."
D · "100% of high scorers studied, proving effort is everything."

C — claim, number, caveat.

A is causal (the data is observational); B has no number; D overclaims from a sample of ten. C says exactly what was measured, how big it was, and how much to trust it.

Part C

The Standards

Attack your own story before anyone else can.

You can now build a finding. This part tries to break it: other explanations, other framings, misleading charts, and a real finding that could hurt real companies.

Interrogation 1

Hunt the confounder

Week 8's lesson goes to work: could a third variable (a confounder) drive both study time and passing? Suppose the ten students came from two class sections, A and B.

df["section"] = list("AAABABABBB") # student A..J df.groupby("section")["passed"].mean() section A 0.2 ← 1 of 5 passed B 0.8 ← 4 of 5 passed Name: passed, dtype: float64
  • list("AAABABABBB")one section letter per student, in order A to J
  • ["passed"].mean()True counts as 1, so the mean is the share who passed

Uh-oh

Section B both studies more and passes more. Different instructor? Different exam? Until you can separate section from study time, "study more, pass more" is on shaky ground.

Interrogation 2

Reframe it — does the story survive?

A solid finding holds under a different cut of the same data (it is robust). Instead of comparing mean hours, compare pass rates of high- vs low-study students.

df["high_study"] = df["study_hours"] >= 5 df.groupby("high_study")["passed"].mean() high_study False 0.0 True 1.0 Name: passed, dtype: float64

Two framings, same direction

Mean-gap and pass-rate agree here — good sign. A "finding" that flips when you change the cut was never a finding; it was an artifact of one framing.

Interrogation 3

Charts can lie without a single false number

A Live One

Now do it with a story that could hurt someone

131 contractors. PHP 75.2 billion. And a word in their names.

The Finding

131 firms, 1,263 projects, ₱75.2 billion

Some contractor names in the source export (the file the agency produced) contain the token (marker text) [REVOKED]. Counting them takes one line, and the total is not small.

4.7% of all awarded value

Against a base of 4,841 contractors and ₱1.6 trillion. Large enough to be worth understanding; large enough to do damage if framed wrongly.

rev = df[df.contractor.str.contains( "REVOKED", na=False)] rev.contractor.nunique() 131 len(rev) 1263 # projects round(rev.budget.sum() / 1e9, 1) 75.2 # PHP billions round(rev.budget.sum() / df.budget.sum(), 3) 0.047 # 4.7% of value
  • df.contractor.str.contains("REVOKED", na=False)True for each row whose contractor name contains the text; na=False treats missing names as False
  • df[ … ]keep only those rows: rev is the flagged projects
  • rev.contractor.nunique()how many different firms: 131
Base Rate

4.7% of the money: more or less than you would expect?

A base rate is how common something is overall: the yardstick you compare a group against. The flagged firms are 2.7% of all contractors and hold 3.7% of the projects, but 4.7% of the value.

What it does and does not say

Their projects are larger than average: 3.7% of the projects carry 4.7% of the value. That is a question to report, not evidence of wrongdoing.

If 3 of your 10 friends are left-handed, is that odd? Only if you know that about 1 person in 10 is left-handed. That "1 in 10" is the base rate.

# share of firms round(rev.contractor.nunique() / df.contractor.nunique(), 3) 0.027 # share of projects round(len(rev) / len(df), 3) 0.037 # share of money round(rev.budget.sum() / df.budget.sum(), 3) 0.047
  • revthe 1,263 projects of the 131 flagged firms (previous slide)
  • nunique() / nunique()131 flagged firms out of 4,841 firms; round(…, 3) keeps 3 decimals
  • len(rev) / len(df)1,263 projects out of 34,079
Write The Headline You Want

Then find out whether you are allowed to

✗ "₱75 billion awarded to contractors with revoked licences"

Every number in it is correct. It is still not a claim this data supports, and it names real companies.

What it implies

That the licence was revoked when the project was awarded. Nothing in the dataset says that.

✓ "131 firms holding ₱75.2B in contracts are flagged as licence-revoked in the DPWH export"

States what the data records, and when. A reader can check it and knows its limits.

What changed

The verb. "Awarded to firms with revoked licences" asserts a sequence. "Flagged as revoked in the export" asserts a record.

What The Token Actually Records

Licence status at export time — and that is all

[REVOKED] is part of the contractor name string in the source. It is a snapshot of that firm’s licence when the file was produced.

Three things it does not tell you

When the licence was revoked. Whether it was revoked before or after these awards. Whether any of these particular projects was irregular.

# what the name looks like # ST. TIMOTHY CONSTRUCTION # CORPORATION ([REVOKED] 39196) # the dataset has no revocation DATE. # so this ordering is unavailable: # award_date vs revocation_date # without it, "awarded to revoked # firms" is an inference, not a # measurement
The Three Questions Before You Name Anyone

Applied to this exact finding

Before Publishing

The pre-publication checklist

Verification means checking a claim against evidence before you publish it. Journalists verify sources; data journalists verify pipelines (every step from raw file to final number). Your weeks 3–6 discipline is the verification.

Source the data

Where it came from, who collected it, what it doesn't cover — stated plainly.

Show the work

Reproducible code from raw file to final chart (week 5's rules). "Trust me" is not a methodology.

Seek the counter-story

Ask "what else would explain this?" — and check it — before someone on the internet does it for you.

A Claim The Data Does Support

Concentration — and you can check every number

Set the licence question aside. Here is a finding with no ordering assumption and no accusation in it. Concentration means a few players hold most of something.

48 firms, 27.6% of the money

Out of 4,841 contractors. That is a structural fact about the market, measurable directly, and it implicates nobody of anything.

c = (df.groupby("contractor").budget.sum() .sort_values(ascending=False)) len(c) 4841 round(c.head(48).sum() / c.sum(), 3) 0.276 # top 1% -> 27.6% round(c.head(10).sum() / c.sum(), 3) 0.1 # top 10 -> 10.0%
  • groupby("contractor").budget.sum()total budget per firm
  • .sort_values(ascending=False)biggest firm first
  • c.head(48).sum() / c.sum()the top 48 firms' total as a share of everything: 0.276 = 27.6%
A Scary Number, Divided

Terminated projects: 131, or 0.4%

"131 flood control projects terminated" is true and sounds bad. 0.4% of all projects is the same fact and sounds like a functioning system.

Report both

The count makes it concrete; the rate makes it fair. Giving only one is where the reader gets steered.

(df.status.value_counts(normalize=True) * 100).round(1) status Completed 80.8 On-Going 16.0 For Procurement 2.6 Terminated 0.4 Not Yet Started 0.3 Name: proportion, dtype: float64 # 0.4% = 131 projects. # both numbers, every time.
  • value_counts(normalize=True)the share of projects in each status (shares add to 1)
  • * 100turn shares into percentages: 0.808 becomes 80.8
Source It So It Can Be Rebuilt

Four facts, or the claim is unverifiable

A source is where your data came from: who produced it, and where you got it. A reader who cannot reproduce your number has to take your word. Four short facts remove that. (CC0 means the data is free to reuse; a mirror is a copy hosted elsewhere.)

Put it under the chart

Not in an appendix. The person who needs it is looking at the figure, not reading your methodology section.

# the caption that makes it checkable Source: DPWH flood control projects, 2016-2026, via BetterGov.ph mirror (CC0). n = 34,079 contracts. Budget summed by contractor; figures in PHP as recorded. Accessed 2026-09-20. # dataset, scope, n, method, date. # five lines. every published figure.
The Pre-Publication Checklist

Everything weeks 4–9 taught, in one pass

CheckFromThe question
Column means what I thinkweek 4Have I looked at .unique() on the field my claim rests on?
I know what I droppedweek 5Row count before and after — and which rows went?
Derived values are possibleweek 6Min, max, and how many are impossible?
I know my denominatorweek 7Does count match len(df)? If not, say so.
The relationship survives splittingweek 8Does it hold within regions, or only on average?
The verb matches the evidenceweek 9Am I asserting a sequence the data cannot order?
A stranger can rebuild itweek 9Source, scope, n, method, date — under the chart.
The Story, Written Out

Finding, number, method, limit

PartThe sentence
FindingFlood control contracting is concentrated: 1% of contractors hold over a quarter of the awarded value.
Number48 of 4,841 firms account for 27.6% of ₱1.6 trillion; the largest single firm holds ₱28.85B across 242 projects.
MethodSummed budget by contractor over 34,079 DPWH projects, 2016–2026, from the BetterGov mirror (CC0).
LimitConcentration is not evidence of wrongdoing. Large firms win large contracts; this measures structure, not conduct.
What This Dataset Cannot Answer

Know the ceiling before you start climbing

QuestionWhy the data cannot reach it
Was this project good value?No quality, inspection or outcome field exists. Budget is not worth.
Was money misused?amount_paid is unpopulated. There are no disbursement records here at all.
Did flooding actually decrease?No flood-incidence data is joined. The dataset describes spending, not effect.
Was a licence revoked before the award?No revocation date. The token is a snapshot at export time.
Who decided each award?No procurement or approval fields. The decision process is not recorded.
Then Go And Ask Someone

The data produces the question; a human answers it

"Region VI’s median project takes 93 days longer than Region I’s" is not a story yet. It is a good question to put to the agency, an engineer, or a contractor.

This is the step that makes it journalism

Analysis stops at the pattern. Reporting takes the pattern to someone who knows why, and publishes their answer alongside it — including when they decline to comment.

# what you bring to the interview - the number (93 days) - the method (median duration by region, n=881 and n=2,730) - the date range (2016-2026) - what you are NOT claiming ("this is not an accusation of delay or mismanagement") # a precise question gets a real answer. # a vague one gets a press release.
Publish A Correction Policy

Being right is a process, not a moment

You will get something wrong. The difference between a credible outlet and a discredited one is entirely what happens next.

Say where the data and code live

A link to the notebook is the strongest possible signal that you expect to be checked — and it is how errors get found while they are still small.

# publish alongside every data story Data: DPWH flood control projects, 2016-2026, BetterGov mirror (CC0) Code: github.com/.../analysis.ipynb Contact: <an actual address> Corrections: any change to a figure is logged at the foot of this page with the date and what changed. # five lines. it is the difference # between a claim and journalism.
The Limit Is Part Of The Finding

Not a disclaimer bolted on at the end

A caveat written last reads as a retreat. A caveat written into the claim reads as precision — and it is what stops a careful reader from dismissing the whole piece.

Hedging vs accuracy

"May possibly suggest" is hedging. "Measures structure, not conduct" is accuracy. The first weakens a claim; the second defines it.

# weak — sounds unsure of itself "This might possibly indicate some kind of irregularity, perhaps." # strong — sure, and precise "48 firms hold 27.6% of awarded value. Concentration is not evidence of wrongdoing; it is a question worth asking the agency."
Your Turn · 8 min

Take a finding all the way

1 · Find something

Any pattern in the flood data: by region, by year, by contractor, by duration.

2 · Write the four parts

Finding, number, method, limit. One sentence each — the table above is the template.

3 · Attack it

Swap with a neighbour. Their job: find the assumption your sentence smuggles in. Every claim has one.

4 · Rewrite

Fix the verb, not the number. Most bad claims are a correct figure attached to an overreaching verb.

Quick Check

Tap to reveal

You find 131 contractors with "[REVOKED]" in their name holding ₱75.2B in projects. Which headline does this dataset support?

A · "₱75 billion awarded to firms with revoked licences"
B · "131 firms holding ₱75.2B in contracts are flagged licence-revoked in the DPWH export"
C · "Government ignored revoked licences worth ₱75 billion"
D · "4.7% of flood control spending went to unlicensed contractors"
B. The token records licence status at export time, inside the name string. There is no revocation date, so the ordering that A, C and D all assume — revoked first, awarded second — cannot be checked. B states what was recorded and when. Same numbers; only B is a measurement rather than an accusation.
Quick Check

Tap to reveal

Which of the five near-misses from weeks 4–8 would have been the most damaging to publish, and why?

A · The −17 day duration — it is obviously impossible
B · The ₱1.17T "completed but never paid" — it was the most exciting, so least likely to be checked, and it names real firms
C · The latitude/longitude correlation — the strongest coefficient
D · The 202-day median — it affects the most rows
B. The damage of a false claim scales with how eagerly it is shared and who it names. The ₱1.17T figure was a trillion-peso accusation about real companies, resting on a column that was simply never populated. A was self-evidently a data error; C nobody would publish; D is a caveat, not an accusation.
Quick Check

Tap to reveal

Draft headline: "Studying CAUSES passing, data shows." Based on the lab's 10-student comparison. What's wrong with it?

A · Nothing — the gap was large
B · Causal claim from observational data, tiny n, and an unchecked section confounder
C · Only that it should say "students"
D · Headlines can't contain the word "data"

B — three strikes in seven words.

Nobody randomized study hours (no causal license), ten students can't carry "shows," and Part C found section B passing at 4× section A's rate. The honest headline: "Passing students studied more — small sample, cause unclear."

Glossary

Every new word from today, in one line each

Finding the story

data journalism
reporting news whose evidence is a dataset
story
a finding a reader should know, told so they can check it
angle
the particular question a story takes on its data
comparison
a number set next to another group, time or expectation
context
surrounding facts that show if a number is big or normal
denominator
what you divide by: per year, per project, per person
per-capita
per person: count ÷ population
base rate
how common something is overall; the yardstick
caveat
the honest limit attached to a claim
outcome / driver
what you explain / what you suspect explains it

Checking the story

source
who produced the data and where you got it
verification
checking a claim against evidence before publishing
unpopulated column
exists but never filled in; amount_paid is 0 on every row
confounder
a third variable that drives both things you compare
robust
still holds when you slice the data another way
truncated axis
bars not starting at zero: exaggerates gaps
cherry-picking
choosing the window or cases that suit the claim
concentration
a few players holding most of something
export / token
the file an agency produced / marker text inside a field
correction policy
a public promise to log and fix errors
observational data
recorded as it happened; shows association, not cause
defamation
a false published claim that damages a reputation
This Week's Lab

Build a finding you can defend

Ten students, a pass/fail outcome, and a suspected driver. You'll make the comparison, chart it, and write the one-sentence finding — number and caveat included. ~45 minutes.

You'll practise

Turning a score into a pass flag, groupby() comparisons with a size and an n, a zero-based bar chart, and the claim–number–caveat sentence.

Stretch, if you're quick

Rates versus counts, and a check for a third variable (tutoring) that could explain the gap.

Recap

Four things to carry out

Readings

Before next week

Next Week

Data Storytelling

Narrative, dashboards and communication — carrying your finding to an audience that has thirty seconds and no context.

DS 227 · Knowledge Discovery in Data