One dataset, the whole toolkit: load it, clean it, analyse it, visualise it — then tell the class the story you found.
Programming for Data Science · University of the Philippines Cebu
The full arc from raw data to a finding — each stage a skill you have.
A project that's reproducible, documented, and honest.
Tell the story clearly, and stand behind it in Q&A.
No new syntax today. This week is about putting eleven weeks together into something you can show — and defend.
A clear plan for the project and the talk that presents it.
The fixed order of steps that turns raw data into an answer: load, clean, analyse, visualise, communicate.
e.g. one script that runs all five, top to bottom
The one specific thing your project sets out to find, written as a question you can answer with the data.
e.g. "Which region's flood-control projects take longest to finish?"
Something you must hand in. Here: the project repository and the talk.
e.g. a repo with a README, code, tests and pinned requirements
The marking guide: the list of criteria your project is judged on, and what earns full marks for each.
e.g. "It runs: fresh clone, no edits needed"
A short text file at the top of the project saying what it is, how to run it, and where the data comes from.
e.g. README.md (week 11)
A .py file of functions that other code can import (week 4).
e.g. src/clean.py, used as from src.clean import coerce_types
A small function that runs your code on known input and asserts the result is right, automatically.
e.g. def test_coerce_counts_failures(): run by pytest
A tiny, hand-made dataset that a test runs on, with one row for each case you care about.
e.g. five rows: normal, unfinished, no budget, impossible, unparseable
A fact that must stay true if a step worked (week 11).
e.g. "a merge does not change the row count"
Something your data or method cannot show, stated openly next to your finding.
e.g. "the place-name pattern matched only 36% of descriptions"
Python, NumPy, pandas, charts, data sources, text, and reproducibility. Alone, each is a technique. Together, they're a pipeline that answers a question.
You've learned every step in isolation.
The project is where they finally connect.
A pipeline is a kitchen routine: shop (load), wash and chop (clean), cook (analyse), plate (visualise), serve (communicate). You have practised every station; now you cook one whole meal.
From raw data to a finding.
You have used every stage on its own (weeks 5–11). This part puts them in one order, and starts from a question instead of a dataset.
It's rarely linear — you'll loop back to clean more once analysis reveals a problem. But this is the shape of every project.
pd.read_csv)groupby, mergeCleaning routinely takes the most time. That's normal, not a failure.
| Stage | From week | Tools |
|---|---|---|
| Load | 9 | read_csv, SQL, APIs |
| Clean | 3, 4, 10 | files, errors, strings + regex |
| Analyse | 5, 6, 7 | NumPy, pandas, groupby, merge |
| Visualise | 8 | matplotlib, seaborn |
| Communicate | 11, 12 | reproducible repo + this talk |
Nothing here is new. The project is assembly, not more learning — you're wiring together tools you've already used.
Stuck on a stage? Reopen that week's lab — the pattern is there.
Not "here's a dataset" but "does city size predict income?" A clear research question (the one thing your project sets out to find) decides what to load, what to clean, and what counts as an answer.
A good one is specific (names what, where, when), answerable with columns you actually have, and small enough for the time you have.
"I'll explore this CSV and see what's there."
"Which region grew fastest from 2015 to 2024, and why?"
Reproducible, documented, honest.
You know the stages. This part is about packaging them so that someone else can run your project, read it, and trust it.
From the first commit, so nothing is ever lost.
A requirements.txt so it runs on any machine.
What the project asks, and how to run it.
A clear layout — data/ (the data files), src/ (short
for "source": your modules, .py files of functions),
notebooks/ — means a stranger, or future you, can find their way in minutes.
Could someone clone it and reproduce your result? If not, it isn't done.
A notebook needs a person. A script (a .py file you run from the terminal) does not — which means it can be scheduled, tested in CI (automatic checks on every push, week 11), and re-run by someone who has never seen it. Its entry point is the function that starts everything, here main().
Lets the same file be imported by tests and run from the command line. Without it, importing runs your whole pipeline.
def main():the whole pipeline as one function: each line is one stage, written as a function in src/df.to_parquet(PROCESSED_PATH)save the cleaned table; Parquet is a compact file format for tablesif __name__ == "__main__":true only when this file is run directly, not when another file imports it; so main() runs only on purpose$ python -m src.pipelinetyped in a terminal (the $ is the prompt, do not type it): run src/pipeline.py as a moduleWrap what can legitimately fail. Let genuine bugs crash — a bare except converts a bug into a wrong number.
Handle the exception you predicted; assert the invariants you expect. Anything else should stop the run.
except FileNotFoundError:catch only the failure you expected (the data file is missing)raise SystemExit(f"…")stop the whole program with a helpful message that points to the READMEassert "budget" in df.columnscheck the file has the column the rest of the code needs; the schema is the list of columns and their typesexcept: passa bare except: catches every error, including real bugs, and carries on as if nothing happenedTwo sentences. What question does this answer, and on what data?
The exact commands, in order, starting from a fresh clone. Test them on someone else’s machine.
Source, licence, date accessed — and how to obtain it, since data/ is gitignored.
The README (README.md, a text file at the top of the repo) is the first thing a marker opens. Write it first, while you still know what is obvious and what is not. A README written at the end is written by someone who has forgotten what confused them on day one — namely, you.
Between the code cells, write markdown that explains the why: the question, each choice, and what a result means. A reader should follow without running anything. (A markdown cell is a notebook cell of formatted text, not code.)
A reader has to reverse-engineer your intent.
Each step is introduced, justified, and interpreted.
The analysis is the same. What changes is whether anyone else can run it.
Every directory here answers a question someone will have. The structure is the documentation you do not have to write.
If a cleaning step is wrong you rerun it. If you edited the raw file by hand, that is gone.
data/raw/, data/processed/the files as downloaded (never edited) and the files your code writessrc/__init__.pyan empty file that turns src/ into a package, so import src.clean workstests/test_clean.pythe tests for clean.py; pytest finds files named test_…notebooks/01-explore.ipynbexploring and the story; the reusable logic lives in src/A cell you scroll back to and re-run is a function that has not been written yet. Extracting it (moving the code into a named function in src/) is what makes it testable.
If the name needs an "and", it is two functions. coerce_types, not clean_and_derive_and_filter.
out = df.copy()work on a copy, so the table passed in stays unchangedpd.to_numeric(…, errors="coerce")turn text into numbers; anything unreadable, like "n/a", becomes NaN instead of an errorpd.to_datetime(…)the same for dates, for both date columns in turnreturn outhand the cleaned copy back to whoever called the functionA docstring is the text in triple quotes right under def (week 4). A docstring repeating the function name is noise. One explaining a decision is the only place that reasoning survives.
You will not remember why errors="coerce" was deliberate. The docstring is where that goes.
df.dropna(subset=subset)drop a row only if one of the listed columns is missing, not any columnreturn out, before - len(out)hand back two things: the smaller table and how many rows were dropped, so you can report itEvery row you dropped is a decision. Say which, and why.
Small sample? Missing data? Name it before someone else does. A limitation is something your data or method cannot show.
Report the result you found, not the one you wanted. (To cherry-pick is to show only the results that suit you.)
The thread from DS 227 (the companion course), one last time: every default is a choice. Making yours visible is what makes the work trustworthy.
Honest and modest beats impressive and unrepeatable.
A test should fail for one reason, and be readable when it does.
Do not test against the real dataset. Write the five rows that represent every case you care about, including the broken ones you found this term. That tiny hand-made table is a fixture.
Normal, unfinished, no budget, the impossible row (it ends before it starts; the flood data has one, week 11), and an unparseable value. Every one is a bug this class actually met.
pd.DataFrame([{…}, {…}])build a table from a list of dictionaries: one dictionary per row, keys are column names... and start "05-23"slide shorthand for "the other columns"; the real fixture spells every column out.replace("", pd.NA)turn empty text into a proper missing value, as read_csv wouldA test is a function that runs your code on known input and asserts the result. Assert the invariant, the property that must hold if the step worked. Those survive refactoring (rewriting how code works without changing what it does); assertions about internals do not.
Row counts, impossible values, NaN counts: you already run them by hand. A test runs them on every change, including the ones you make at 2am.
def test_…():pytest runs every function whose name starts with test_; it passes if no assert failsout.budget.isna().sum() == 2exactly two budgets became NaN: the empty one (A3) and "n/a" (A5)(out.duration < 0).sum() == 1exactly one row finishes before it starts, and the next line checks it is A4cannot regressto regress is for a fixed bug to come back; this test would catch it if it didReading it: pytest tests/ -q runs every test in tests/ quietly; each . is one test that passed (a failure shows F); 5 passed in 7.83s is the summary. Five dots. That is the whole return on the discipline: the checks you have been running by hand all term now run automatically, and a change that breaks one tells you immediately rather than in a chart three weeks later.
And before you hand anything in. A failing test the night before is a gift; the same bug found by your marker is not.
You are not aiming for coverage (in testing, the share of your code lines that some test runs). You are encoding the mistakes that are expensive when they recur.
Assume the library works. Test your logic — the decisions you made about this data.
shape invariantsfacts about rows and columns: counts that must not changedomain rulesfacts about the real world the data describes: a project cannot end before it startsuse approx0.1 + 0.2 is not exactly 0.3 in a computer; pytest.approx allows a tiny differenceIn the pipeline, which stage usually takes the most time?
B — Clean.
Real data is messy: missing values, wrong types, stray text. Cleaning routinely dominates a project. Budget for it, and document what you did.
Expect cleaning to be the long pole. Planning for that is half the battle.
Most of the work is getting the data clean.
Back to turn the project into a talk.
Tell the story; stand behind it.
Your project runs and can be trusted. This part is about explaining it to a room in a few minutes, and answering their questions honestly.
If the audience remembers a single sentence, what is it? Decide that first, then build every slide to carry them toward it. Everything else is support.
That sentence is your finding: the answer to your research question, in plain words, with a number if you can.
"Here is everything I did, in order."
"Visayan cities grew fastest — here's the evidence."
| Beat | Answers |
|---|---|
| Question | What did you set out to learn? |
| Data | Where did it come from, and how clean was it? |
| Method | What did you actually do to it? |
| Finding | What's the answer — the one message? |
| Limitation | Where should we not trust it? |
This arc fits five minutes or fifty. Each beat is one or two slides — no more.
If a slide doesn't serve one of these beats, cut it.
A slide makes a single point. Split it if it makes two.
A picture lands from the back row; a grid of numbers doesn't.
Title, axes, units — and honest baselines, as in Week 8.
These are the same principles this very deck follows: little text, one concept, a visual that carries the point.
Readable from the back, understood in five seconds.
Nobody wants a line-by-line tour. Explain what you asked, what you found, and what it means. Rehearse aloud once — it exposes every rough transition.
Speak to the finding; let charts do the arguing.
Read your slides, or walk through every cell.
| Stage | Week | What you use |
|---|---|---|
| Get it | 9 | read_csv / read_sql / requests — and cache it |
| Shape it | 6–7 | dtypes, loc, groupby, merge, transform |
| Compute it | 5 | NumPy: vectorise, mask, axis — never a loop |
| Parse the text | 10 | .str, regex — then report coverage |
| Show it | 8 | subplots, labels, bars from zero, saved with bbox_inches |
| Ship it | 11 | git, pinned requirements, tests, README |
Six stages: the same five-stage pipeline, with Clean split into "shape it", "compute it" and "parse the text", and Communicate called "ship it". You have done every one of them separately. The project is the first time they are one artifact — and the ordering is not arbitrary: each stage assumes the one above it was done properly.
Nobody wants to read your for loop on a projector. Show the one function where you made a judgement call, and explain the call. (Boilerplate is the routine code every project repeats, like imports and axis labels.)
Lead with the output. Have the notebook open in another tab for the question you hope someone asks.
README.md: the question, how to run it, where the data comes fromrequirements.txt with pinned versionssrc/: the cleaning and analysis as functions with docstringsnotebooks/: the story, text between the code, runs top to bottomtests/: at least one shape invariant, one domain rule, one bug you hitdata/raw/ untouched; data and secrets not committedA deliverable is something you must hand in. Everything on this slide was covered today; the next two slides say how it is marked. Check the course page for the due date, talk length and where to submit.
A classmate clones the repo, installs requirements.txt in a fresh environment, runs it top to bottom and gets your numbers, with no edits.
| Criterion | What earns full marks |
|---|---|
| It runs | Fresh clone, fresh environment, pinned requirements, Restart & Run All — no edits needed |
| Structure | Logic in src/, notebooks for exploring; raw data never modified |
| Correctness | Tests exist and pass; invariants asserted at the steps that can silently break |
| Readability | Functions do one job; docstrings say why; names mean something |
| Honesty | Coverage and row counts reported; no bare excepts; limits stated |
"It runs" is first because everything else is unverifiable without it. A brilliant analysis nobody can execute scores below a modest one that works on the first try.
| Criterion | In plain words | Check it yourself by… |
|---|---|---|
| It runs | Someone else can run it and get your numbers | Clone to a new folder, new environment, pip install -r requirements.txt, Restart & Run All |
| Structure | Reusable code lives in files of functions, not in notebook cells | Can the notebook import your cleaning from src/? Is data/raw/ as downloaded? |
| Correctness | You proved the steps that could quietly go wrong | pytest -q shows only dots; an assert after each merge and cleaning step |
| Readability | A stranger can follow it | Every function name says its one job; docstrings say why; no df2, x |
| Honesty | You show what you dropped and what you cannot claim | Rows before/after cleaning reported; regex coverage reported; limitations in the README and talk |
If you can tick the right-hand column for every row, you know before you submit how the project will score. The Your Turn activity and the lab both practise these checks.
You will read this code far more often than you write it — and mostly when something is wrong and you are in a hurry.
One frame in scope, call it df. Two, and they both need real names — projects and regions.
1e8scientific notation for 100,000,000 (a 1 followed by 8 zeros): here, budgets over ₱100 milliondf.b, "r"one-letter names force the reader to go and look up what they holdby_region = …the name says what is inside: the average budget of large projects, per regionIf the name needs "and", split it. If it is longer than a screen, it is doing several things.
Paths relative, README says where the file comes from, data/raw/ untouched.
Bare except, an extraction with unreported coverage, a merge with no row-count check.
Then the three that catch the rest: are the requirements pinned, do the tests pass, and does it run from a fresh clone? Every one of these has cost this course real time in a previous week.
Everything below assumes you can load, shape, compute and ship data reliably. That assumption is what you spent eleven weeks earning.
SQL and testing pay off immediately and in every job. New libraries are easy once the habits are in place.
None of these is needed for the project. They are names to look up later: polars and duckdb are faster table tools, scikit-learn is for machine learning, Flask/FastAPI build web APIs, streamlit turns a script into a web dashboard. Window functions and CTEs are more advanced SQL; parametrize runs one test on many inputs.
pandas 3 changed the copy rules under you mid-course. It will happen again.
Raw data read-only, logic in modules, tests on invariants, pinned dependencies, a README. That layout has outlasted every library in it.
Most people who can analyse data cannot hand it over. Being the person whose code runs on someone else’s machine is rarer than it should be.
Eleven weeks ago "it works on my machine" sounded like it was finished. Knowing why that is the beginning of the work, not the end of it, is the course.
Pick the two cells you have re-run most. Move them into src/clean.py with docstrings that say why.
Five hand-written rows covering: normal, missing, and one broken case you actually hit this term.
One shape invariant, one domain rule, one for the bug. Run pytest -q in a terminal, from the project folder.
Change a function so a test fails. Read the failure output — that is what it will look like at 2am.
Step 4 matters: a test you have never seen fail is a test you do not know the meaning of.
Your clean() function works in the notebook. What is the first thing to do before relying on it?
Which of these is worth writing a test for?
pd.to_numeric converts "123" to 123pytest.approx.You know your project's weak spots better than anyone — so name them first. When a question goes past what you tested, say so plainly; guessing is what loses trust.
List the three questions you'd least like to be asked, and answer them.
"That's outside what I tested" beats a confident wrong answer.
What's the best thing to decide first when building your presentation?
B — the single message.
Decide the one sentence the audience should leave with, then build every slide toward it. Design choices come after the message, never before.
Message first, slides second, styling last. That order keeps a talk focused.
One message drives the whole deck.
| Check | Ask yourself |
|---|---|
| Reproducible | Could someone clone it and rerun it cleanly? |
| Documented | Does the README + notebook explain the why? |
| Honest | Are cleaning choices and limitations stated? |
| Clear | Is there one message a stranger would grasp? |
Four yeses and you're ready. These are the marks of professional data work — not just a passing project.
Run it once from a fresh clone before you call it done.
Every project is the same arc, and you own each stage.
Reproducible, documented, honest — from the first commit.
One message, clear slides, and integrity in Q&A.
One sentence: you can take a question and a raw dataset and turn them into an answer other people can trust.
Everything after this is practice and depth.
.py file of functions you can import.py file run from the terminal / its starting function, main()__main__ guardif __name__ == "__main__": runs only when the file is run directlyexcept: with no error named; hides real bugsdef; says whyDS 227 uses this toolkit for the full data-science method — modelling, evaluation, and more.
Find a dataset you care about and run the whole pipeline on it, start to finish.
The syntax fades if unused; the workflow stays. Keep the habits — Git, clean data, honest charts — and the tools will always come back.
Build things. Nothing teaches like a project you actually finish.
You came in unsure you could program. You're leaving able to load, clean, analyse, visualise, and present real data. Well done.
DS 208 · Programming for Data Science · University of the Philippines Cebu