DS 208 · Week 12

Integrative Project & Presentations

One dataset, the whole toolkit: load it, clean it, analyse it, visualise it — then tell the class the story you found.

Programming for Data Science · University of the Philippines Cebu

Session Map

Everything, on one project

Words for Today

Ten words for the project

pipeline

The fixed order of steps that turns raw data into an answer: load, clean, analyse, visualise, communicate.

e.g. one script that runs all five, top to bottom

research question

The one specific thing your project sets out to find, written as a question you can answer with the data.

e.g. "Which region's flood-control projects take longest to finish?"

deliverable

Something you must hand in. Here: the project repository and the talk.

e.g. a repo with a README, code, tests and pinned requirements

rubric

The marking guide: the list of criteria your project is judged on, and what earns full marks for each.

e.g. "It runs: fresh clone, no edits needed"

README

A short text file at the top of the project saying what it is, how to run it, and where the data comes from.

e.g. README.md (week 11)

module

A .py file of functions that other code can import (week 4).

e.g. src/clean.py, used as from src.clean import coerce_types

test

A small function that runs your code on known input and asserts the result is right, automatically.

e.g. def test_coerce_counts_failures(): run by pytest

fixture

A tiny, hand-made dataset that a test runs on, with one row for each case you care about.

e.g. five rows: normal, unfinished, no budget, impossible, unparseable

invariant

A fact that must stay true if a step worked (week 11).

e.g. "a merge does not change the row count"

limitation

Something your data or method cannot show, stated openly next to your finding.

e.g. "the place-name pattern matched only 36% of descriptions"

Where We've Been

Eleven weeks, one workflow

Python, NumPy, pandas, charts, data sources, text, and reproducibility. Alone, each is a technique. Together, they're a pipeline that answers a question.

The parts

You've learned every step in isolation.

The whole

The project is where they finally connect.

A pipeline is a kitchen routine: shop (load), wash and chop (clean), cook (analyse), plate (visualise), serve (communicate). You have practised every station; now you cook one whole meal.

Part A

The Pipeline

From raw data to a finding.

You have used every stage on its own (weeks 5–11). This part puts them in one order, and starts from a question instead of a dataset.

The Mental Model

Five stages, start to finish

Load Clean Analyse Visualise Communicate most of the effort lives in Clean — expect that
You Already Have Each Piece

Map the stages to your weeks

StageFrom weekTools
Load9read_csv, SQL, APIs
Clean3, 4, 10files, errors, strings + regex
Analyse5, 6, 7NumPy, pandas, groupby, merge
Visualise8matplotlib, seaborn
Communicate11, 12reproducible repo + this talk
Start Here

A project begins with a question

Not "here's a dataset" but "does city size predict income?" A clear research question (the one thing your project sets out to find) decides what to load, what to clean, and what counts as an answer.

A good one is specific (names what, where, when), answerable with columns you actually have, and small enough for the time you have.

Weak start

"I'll explore this CSV and see what's there."

Strong start

"Which region grew fastest from 2015 to 2024, and why?"

Part B

Structuring the Project

Reproducible, documented, honest.

You know the stages. This part is about packaging them so that someone else can run your project, read it, and trust it.

Set It Up Right

Four things every project needs

An Entry Point

So it runs without a human clicking cells

A notebook needs a person. A script (a .py file you run from the terminal) does not — which means it can be scheduled, tested in CI (automatic checks on every push, week 11), and re-run by someone who has never seen it. Its entry point is the function that starts everything, here main().

The __main__ guard

Lets the same file be imported by tests and run from the command line. Without it, importing runs your whole pipeline.

# src/pipeline.py def main(): df = load_raw(RAW_PATH) df = coerce_types(df) df = add_duration(df) df, dropped = drop_unusable(df, ["budget"]) log.info("dropped %s rows", dropped) df.to_parquet(PROCESSED_PATH) return df if __name__ == "__main__": main() $ python -m src.pipeline
  • def main():the whole pipeline as one function: each line is one stage, written as a function in src/
  • df.to_parquet(PROCESSED_PATH)save the cleaned table; Parquet is a compact file format for tables
  • if __name__ == "__main__":true only when this file is run directly, not when another file imports it; so main() runs only on purpose
  • $ python -m src.pipelinetyped in a terminal (the $ is the prompt, do not type it): run src/pipeline.py as a module
Fail Loudly, In The Right Place

A pipeline that dies early is easier than one that lies

Wrap what can legitimately fail. Let genuine bugs crash — a bare except converts a bug into a wrong number.

Catch narrow, assert wide

Handle the exception you predicted; assert the invariants you expect. Anything else should stop the run.

try: df = pd.read_csv(path) except FileNotFoundError: raise SystemExit(f"missing {path} — see README") assert len(df) > 0, "loaded an empty file" assert "budget" in df.columns, "schema changed" # NEVER: # try: ... # except: pass # that is how a wrong answer ships
  • except FileNotFoundError:catch only the failure you expected (the data file is missing)
  • raise SystemExit(f"…")stop the whole program with a helpful message that points to the README
  • assert "budget" in df.columnscheck the file has the column the rest of the code needs; the schema is the list of columns and their types
  • except: passa bare except: catches every error, including real bugs, and carries on as if nothing happened
The README Is Part Of The Code

Three sections, written on day one

Code Plus Story

A notebook should read top to bottom

Between the code cells, write markdown that explains the why: the question, each choice, and what a result means. A reader should follow without running anything. (A markdown cell is a notebook cell of formatted text, not code.)

Just code

A reader has to reverse-engineer your intent.

Code + narrative

Each step is introduced, justified, and interpreted.

A Project Is Not A Notebook

It is a notebook plus the things around it

The analysis is the same. What changes is whether anyone else can run it.

The Layout

Where a stranger would look

Every directory here answers a question someone will have. The structure is the documentation you do not have to write.

raw/ is read-only, always

If a cleaning step is wrong you rerun it. If you edited the raw file by hand, that is gone.

flood_project/ data/ raw/ # never edited processed/ # written by code src/ __init__.py load.py # reading files/APIs clean.py # one job per module analyse.py tests/ test_clean.py notebooks/ 01-explore.ipynb README.md requirements.txt
  • data/raw/, data/processed/the files as downloaded (never edited) and the files your code writes
  • src/__init__.pyan empty file that turns src/ into a package, so import src.clean works
  • tests/test_clean.pythe tests for clean.py; pytest finds files named test_…
  • notebooks/01-explore.ipynbexploring and the story; the reusable logic lives in src/
Cells Become Functions

The moment you need it twice

A cell you scroll back to and re-run is a function that has not been written yet. Extracting it (moving the code into a named function in src/) is what makes it testable.

One job per function

If the name needs an "and", it is two functions. coerce_types, not clean_and_derive_and_filter.

# src/clean.py def coerce_types(df): out = df.copy() out["budget"] = pd.to_numeric( out.budget, errors="coerce") for col in ("start_date", "completion_date"): out[col] = pd.to_datetime(out[col], errors="coerce") return out # .copy() so the caller's frame is # never modified by surprise
  • out = df.copy()work on a copy, so the table passed in stays unchanged
  • pd.to_numeric(…, errors="coerce")turn text into numbers; anything unreadable, like "n/a", becomes NaN instead of an error
  • pd.to_datetime(…)the same for dates, for both date columns in turn
  • return outhand the cleaned copy back to whoever called the function
Docstrings Say Why

The code already says what

A docstring is the text in triple quotes right under def (week 4). A docstring repeating the function name is noise. One explaining a decision is the only place that reasoning survives.

Write it for yourself in March

You will not remember why errors="coerce" was deliberate. The docstring is where that goes.

def drop_unusable(df, subset): """Drop rows missing any of `subset`, and report how many went. Deliberately NOT a bare df.dropna(): that removes rows for reasons unrelated to the question and silently changes which population the analysis is about. """ before = len(df) out = df.dropna(subset=subset) return out, before - len(out)
  • df.dropna(subset=subset)drop a row only if one of the listed columns is missing, not any column
  • return out, before - len(out)hand back two things: the smaller table and how many rows were dropped, so you can report it
Integrity Over Polish

Show the choices you made

Now You Can Test It

Five rows, not 34,079

A test should fail for one reason, and be readable when it does.

Build The Fixture By Hand

Each row exists to exercise one case

Do not test against the real dataset. Write the five rows that represent every case you care about, including the broken ones you found this term. That tiny hand-made table is a fixture.

These five are the course

Normal, unfinished, no budget, the impossible row (it ends before it starts; the flood data has one, week 11), and an unparseable value. Every one is a bug this class actually met.

RAW = pd.DataFrame([ {"id":"A1", "budget":"1000000", ...}, # normal {"id":"A2", ... "completion_date":""}, # unfinished {"id":"A3", "budget":"", ...}, # no budget {"id":"A4", start "05-23", end "05-06"}, # impossible {"id":"A5", "budget":"n/a", ...}, # unparseable ]).replace("", pd.NA) # .replace("", pd.NA) because read_csv # would have — and dropna only sees NaN
  • pd.DataFrame([{…}, {…}])build a table from a list of dictionaries: one dictionary per row, keys are column names
  • ... and start "05-23"slide shorthand for "the other columns"; the real fixture spells every column out
  • .replace("", pd.NA)turn empty text into a proper missing value, as read_csv would
Test The Invariant

Not the implementation

A test is a function that runs your code on known input and asserts the result. Assert the invariant, the property that must hold if the step worked. Those survive refactoring (rewriting how code works without changing what it does); assertions about internals do not.

These are your checks, executable

Row counts, impossible values, NaN counts: you already run them by hand. A test runs them on every change, including the ones you make at 2am.

def test_coerce_counts_failures(): out = coerce_types(RAW) assert out.budget.isna().sum() == 2 # A3 empty + A5 "n/a" — neither raised def test_surfaces_the_impossible_row(): out = add_duration(coerce_types(RAW)) assert (out.duration < 0).sum() == 1 assert out.loc[out.duration < 0, "id"].iloc[0] == "A4" # the impossible row, encoded so it cannot regress
  • def test_…():pytest runs every function whose name starts with test_; it passes if no assert fails
  • out.budget.isna().sum() == 2exactly two budgets became NaN: the empty one (A3) and "n/a" (A5)
  • (out.duration < 0).sum() == 1exactly one row finishes before it starts, and the next line checks it is A4
  • cannot regressto regress is for a fixed bug to come back; this test would catch it if it did
It Runs In Under Eight Seconds

Every time, instead of when you remember

$ pytest tests/ -q ..... [100%] 5 passed in 7.83s
What Is Worth Testing

Three categories, and skip the rest

You are not aiming for coverage (in testing, the share of your code lines that some test runs). You are encoding the mistakes that are expensive when they recur.

Do not test pandas

Assume the library works. Test your logic — the decisions you made about this data.

TEST: shape invariants "a merge must not change row count" domain rules "duration cannot be negative" the bugs you already hit "unparseable budget -> NaN, not a crash" DO NOT TEST: that pandas can add numbers exact float equality (use approx) plotting output
  • shape invariantsfacts about rows and columns: counts that must not change
  • domain rulesfacts about the real world the data describes: a project cannot end before it starts
  • use approx0.1 + 0.2 is not exactly 0.3 in a computer; pytest.approx allows a tiny difference
Quick Check

Tap to reveal

In the pipeline, which stage usually takes the most time?

A · Load
B · Clean
C · Visualise
D · Communicate

B — Clean.

Real data is messy: missing values, wrong types, stray text. Cleaning routinely dominates a project. Budget for it, and document what you did.

Break

  Five minutes

Back to turn the project into a talk.

Part C

Presenting

Tell the story; stand behind it.

Your project runs and can be trusted. This part is about explaining it to a room in a few minutes, and answering their questions honestly.

Lead With The Finding

One talk, one message

If the audience remembers a single sentence, what is it? Decide that first, then build every slide to carry them toward it. Everything else is support.

That sentence is your finding: the answer to your research question, in plain words, with a number if you can.

Weak

"Here is everything I did, in order."

Strong

"Visayan cities grew fastest — here's the evidence."

A Shape That Works

Five beats, in order

BeatAnswers
QuestionWhat did you set out to learn?
DataWhere did it come from, and how clean was it?
MethodWhat did you actually do to it?
FindingWhat's the answer — the one message?
LimitationWhere should we not trust it?
Design For The Room

Slides support you — they aren't the report

On The Day

Tell the story, not the code

Nobody wants a line-by-line tour. Explain what you asked, what you found, and what it means. Rehearse aloud once — it exposes every rough transition.

Do

Speak to the finding; let charts do the arguing.

Don't

Read your slides, or walk through every cell.

Eleven Weeks, One Pipeline

Every tool in the place it belongs

StageWeekWhat you use
Get it9read_csv / read_sql / requests — and cache it
Shape it6–7dtypes, loc, groupby, merge, transform
Compute it5NumPy: vectorise, mask, axis — never a loop
Parse the text10.str, regex — then report coverage
Show it8subplots, labels, bars from zero, saved with bbox_inches
Ship it11git, pinned requirements, tests, README
Presenting Code

Show the decision, not the syntax

Nobody wants to read your for loop on a projector. Show the one function where you made a judgement call, and explain the call. (Boilerplate is the routine code every project repeats, like imports and axis labels.)

Results first, code on request

Lead with the output. Have the notebook open in another tab for the question you hope someone asks.

DO show: the function with the decision in it (5-10 lines, with its docstring) the test that encodes an invariant ("a merge must not change row count") the output, large enough to read DON'T show: imports boilerplate plotting calls anything you would scroll
The Deliverables

What you hand in: a repository and a talk

1 · The repository (your code, on Git)

  • README.md: the question, how to run it, where the data comes from
  • requirements.txt with pinned versions
  • src/: the cleaning and analysis as functions with docstrings
  • notebooks/: the story, text between the code, runs top to bottom
  • tests/: at least one shape invariant, one domain rule, one bug you hit
  • data/raw/ untouched; data and secrets not committed

2 · The talk (to the class)

  • Five beats: question, data, method, finding, limitation
  • One message the audience can repeat afterwards
  • Labelled charts, not tables of numbers or scrolling code
  • One function that holds a decision, shown and explained
  • Q&A: name your weak spots before you are asked
How The Project Is Marked

Runs-on-another-machine is weighted highest

CriterionWhat earns full marks
It runsFresh clone, fresh environment, pinned requirements, Restart & Run All — no edits needed
StructureLogic in src/, notebooks for exploring; raw data never modified
CorrectnessTests exist and pass; invariants asserted at the steps that can silently break
ReadabilityFunctions do one job; docstrings say why; names mean something
HonestyCoverage and row counts reported; no bare excepts; limits stated
The Rubric, In Plain Words

Each criterion is a check you can run yourself

CriterionIn plain wordsCheck it yourself by…
It runsSomeone else can run it and get your numbersClone to a new folder, new environment, pip install -r requirements.txt, Restart & Run All
StructureReusable code lives in files of functions, not in notebook cellsCan the notebook import your cleaning from src/? Is data/raw/ as downloaded?
CorrectnessYou proved the steps that could quietly go wrongpytest -q shows only dots; an assert after each merge and cleaning step
ReadabilityA stranger can follow itEvery function name says its one job; docstrings say why; no df2, x
HonestyYou show what you dropped and what you cannot claimRows before/after cleaning reported; regex coverage reported; limitations in the README and talk
Names Are Documentation

The cheapest readability there is

You will read this code far more often than you write it — and mostly when something is wrong and you are in a hurry.

df is fine. df2 is not.

One frame in scope, call it df. Two, and they both need real names — projects and regions.

# meaningless df2 = df[df.b > 1e8] x = df2.groupby("r").mean() # says what it holds large = projects[projects.budget > 1e8] by_region = large.groupby("region").budget.mean() # and the unit, when there is one duration_days = ... budget_php = ...
  • 1e8scientific notation for 100,000,000 (a 1 followed by 8 zeros): here, budgets over ₱100 million
  • df.b, "r"one-letter names force the reader to go and look up what they hold
  • by_region = …the name says what is inside: the average budget of large projects, per region
Read Your Own Code, Once

Six questions, twenty minutes, before you submit

What To Learn Next

This was the toolkit; here is what it unlocks

Everything below assumes you can load, shape, compute and ship data reliably. That assumption is what you spent eleven weeks earning.

Depth before breadth

SQL and testing pay off immediately and in every job. New libraries are easy once the habits are in place.

None of these is needed for the project. They are names to look up later: polars and duckdb are faster table tools, scikit-learn is for machine learning, Flask/FastAPI build web APIs, streamlit turns a script into a web dashboard. Window functions and CTEs are more advanced SQL; parametrize runs one test on many inputs.

immediately useful: SQL beyond SELECT — window functions, CTEs pytest properly — fixtures, parametrize polars / duckdb when pandas stops fitting in memory builds directly on this: scikit-learn (CMSC 173) APIs you write, not just call (Flask/FastAPI) dashboards (streamlit) the habit that carries: extract, test, pin, document
If You Keep One Thing

Working code that someone else can run

Your Turn · 8 min

Refactor one notebook into a project

1 · Extract two functions

Pick the two cells you have re-run most. Move them into src/clean.py with docstrings that say why.

2 · Write the fixture

Five hand-written rows covering: normal, missing, and one broken case you actually hit this term.

3 · Write three tests

One shape invariant, one domain rule, one for the bug. Run pytest -q in a terminal, from the project folder.

4 · Break it on purpose

Change a function so a test fails. Read the failure output — that is what it will look like at 2am.

Quick Check

Tap to reveal

Your clean() function works in the notebook. What is the first thing to do before relying on it?

A · Optimise it — it is called on 34,079 rows
B · Move it into a module and write a test on a handful of hand-made rows
C · Add a try/except so it never crashes
D · Run it on the full dataset again to be sure
B. A notebook cell cannot be tested, imported or reviewed. Extracting it makes all three possible, and a five-row fixture catches the cases the real data may not contain today. C is actively harmful — a bare except turns a bug into a wrong number. D confirms nothing: it already ran, which is why you think it works.
Quick Check

Tap to reveal

Which of these is worth writing a test for?

A · That pd.to_numeric converts "123" to 123
B · That your merge does not change the row count
C · That matplotlib produces a PNG
D · That two floats are exactly equal
B. It is your invariant — a duplicated join key silently doubling rows is a bug this course met in week 7, and the test catches it on every change. A tests pandas, which you may assume works. C tests a library. D is actively bad practice: floating point makes exact equality unreliable, so use pytest.approx.
Standing Behind It

"I don't know" is a strong answer

You know your project's weak spots better than anyone — so name them first. When a question goes past what you tested, say so plainly; guessing is what loses trust.

Prepare

List the three questions you'd least like to be asked, and answer them.

Be honest

"That's outside what I tested" beats a confident wrong answer.

Quick Check

Tap to reveal

What's the best thing to decide first when building your presentation?

A · the colour theme
B · the single message
C · how many slides
D · the font

B — the single message.

Decide the one sentence the audience should leave with, then build every slide toward it. Design choices come after the message, never before.

Before You Submit

The four-point check

CheckAsk yourself
ReproducibleCould someone clone it and rerun it cleanly?
DocumentedDoes the README + notebook explain the why?
HonestAre cleaning choices and limitations stated?
ClearIs there one message a stranger would grasp?
The Course, In One Breath

Load, clean, analyse, show, share

Glossary

This week's words, one line each

Words for today

pipeline
load → clean → analyse → visualise → communicate, in order
research question
the one specific, answerable thing the project sets out to find
deliverable
what you hand in: the repository and the talk
rubric
the marking guide: criteria and what earns full marks
README
text file saying what the project is, how to run it, where data comes from
module
a .py file of functions you can import
test
a function that runs your code on known input and asserts the result
fixture
a tiny hand-made dataset a test runs on
invariant
a fact that must stay true if a step worked
limitation
what your data or method cannot show, stated openly

Also new today

finding
the answer to your question, in one plain sentence
criterion
one thing the rubric judges (plural: criteria)
script / entry point
a .py file run from the terminal / its starting function, main()
__main__ guard
if __name__ == "__main__": runs only when the file is run directly
bare except
except: with no error named; hides real bugs
docstring
text in triple quotes under def; says why
refactor
restructure code without changing what it does
fresh clone
a new copy of the repo on a machine that has never seen it
cherry-pick
show only the results that suit you (don't)
boilerplate
routine code every project repeats
Where To Go Next

This was the foundation

Thank You

Now go build something

You came in unsure you could program. You're leaving able to load, clean, analyse, visualise, and present real data. Well done.

DS 208 · Programming for Data Science · University of the Philippines Cebu