DS 208 · Week 11

Reproducible Workflows

Code someone else can run next year and get the same number — with Git for history, environments for dependencies, and calm debugging.

Programming for Data Science · University of the Philippines Cebu

Session Map

Three habits of professional work

Words for Today · 1 of 2

Eight words for keeping history

reproducibility

Someone else can rerun your work, on their computer, and get the same result.

e.g. a classmate runs your notebook next year and sees the same total

version control

A system that saves named snapshots of your files so you can go back to, compare or share any of them. Git is the one we use.

e.g. "show me the code as it was last Tuesday"

repository (repo)

A project folder that Git is tracking, together with its whole saved history.

e.g. your ds208-project/ folder after git init

commit

One saved snapshot of the project, with a message saying what changed and why.

e.g. git commit -m "Fix region totals"

branch

A separate line of commits where you can try something without touching the working version.

e.g. a branch called feature/regional-breakdown

merge

Bring the commits from one branch into another, combining the two lines of work.

e.g. git merge feature/regional-breakdown while on main

remote

A copy of the repository on another computer, usually a website like GitHub, that everyone shares.

e.g. github.com/you/ds208-project

push / pull

Push sends your new commits up to the remote; pull brings other people's commits down to you.

e.g. git push after committing; git pull before starting

Words for Today · 2 of 2

Nine words for setting up and fixing

dependency

A package your code needs in order to run.

e.g. your analysis depends on pandas and matplotlib

environment

The Python plus the exact set of installed packages that your code runs with. A virtual environment is a private one per project.

e.g. a .venv folder inside your project

package manager (pip / conda)

The tool that downloads and installs packages. pip comes with Python; conda comes with Anaconda.

e.g. pip install pandas

requirements file

A text file, requirements.txt, listing every package the project needs, one per line.

e.g. a line reading pandas==3.0.5

version pin

Writing down the exact version to install, so everyone gets the same one.

e.g. ==3.0.5 in pandas==3.0.5

debugger

A tool that pauses your running program so you can look at every variable and move forward line by line.

e.g. Python's built-in pdb

breakpoint

A marked line where the program pauses and hands control to the debugger.

e.g. writing breakpoint() on the line before the crash

stepping

Running a paused program one line at a time, watching what each line does.

e.g. typing n (next line) at the (Pdb) prompt

assertion

A line that checks something must be true and stops the program with an error if it is not.

e.g. assert len(df) > 0, "no rows loaded"

Where We Left Off

"It works on my machine" isn't enough

You can load, analyse, and plot (weeks 5–10). But if only you can run it, on only your laptop, it isn't science yet. Reproducibility (anyone can rerun it and get the same result) is what makes work trustworthy.

A good recipe lists exact amounts and steps, so anyone can bake the same cake. Today is about writing your analysis like that recipe, not like "a bit of this, a pinch of that".

The risk

Lost history, mismatched versions, bugs you can't retrace.

The fix

Track changes, pin dependencies, debug methodically.

Part A

Git

A save-point machine for your code.

You already save files. This part adds version control: saving named snapshots you can return to, compare, and share with others.

The Motivation

Every change, undoable and explained

Git records snapshots of your project over time. You can undo, see who changed what and why, and merge work from several people without chaos.

A folder Git tracks is a repository (repo); each saved snapshot is a commit.

Where do git commands go?

Into a terminal (the command line): a text window where you type instructions to your computer. Terminal on macOS/Linux, PowerShell or Git Bash on Windows. Not into Python or a notebook cell. Run them from inside your project folder.

Time machine

Roll back to any past snapshot in seconds.

Teamwork

Many people, one project, no overwriting each other.

Git is like the version history in Google Docs, except you decide when a version is saved, and you write a note on each one saying why.

The Everyday Rhythm

Add, commit, push

Stage the files you changed, save a snapshot with a message, and send it to the shared copy. Three commands you'll run dozens of times a day.

To stage a file is to put its changes in the "next snapshot" pile; to push is to upload your commits to the remote, the shared copy (e.g. on GitHub).

git add analysis.py git commit -m "Add city density calc" git push
  • git add analysis.pystage: include this file's current changes in the next commit
  • git commit -m "…"save everything staged as one commit; -m attaches the message in quotes
  • git pushupload your new commits to the remote so others can get them
Reading History

Four commands answer almost every question

Version control earns its keep when something breaks and you need to know what changed. These are how you ask.

--oneline makes log usable

The default format shows four lines per commit. You want one, so you can scan fifty at once.

git log --oneline -10 6b80d5e expand week 10 decks 8e2216e expand week 9 decks 625da22 expand week 8 decks git log --oneline -- path/to/file # one file git show 8e2216e # one commit git diff # unstaged now git diff --staged # what commit will take git blame file.py # who + when, per line
  • git log --oneline -10the last 10 commits, one line each: short ID, then message
  • 6b80d5ea commit's ID (its SHA, a unique fingerprint); type it to point at that commit
  • -- path/to/fileonly the commits that touched that one file
  • git show 8e2216ewhat that one commit changed
  • git diffa diff is the line-by-line changes: here, edits you have not staged yet; --staged shows what the next commit will save
  • git blame file.pyfor each line, which commit (and person) last changed it
The Three Undos

They are not interchangeable — pick by what you want to keep

CommandWhat it undoesHistorySafe to use after pushing?
git restore <file>Uncommitted edits to that fileUntouchedYes — nothing was committed
git revert <sha>A commit, by adding an inverse commitPreserved, moves forwardYes — this is the shared-branch answer
git reset --hard <sha>Everything back to that commitRewrittenNo — it breaks everyone else’s clone
reset Feels Permanent. It Is Not.

reflog remembers what log forgot

A hard reset looks like the commit is gone — it is not in git log any more. But git kept the pointer for weeks.

Verified, not assumed

Run in a scratch repo: after reset --hard HEAD~1 the discarded commit is the second line of git reflog.

git reset --hard HEAD~1 # "second" is gone from git log git reflog 080d423 HEAD@{0}: reset: moving to HEAD~1 1992960 HEAD@{1}: commit: second # get it back git reset --hard 1992960 # reflog is LOCAL and expires (~90d). # it saves you from yourself, not from # a lost laptop
  • HEAD~1HEAD is the commit you are on now; ~1 means "one before it"
  • git reset --hard …move back to that commit and throw away everything after it
  • git reflogGit's private diary of every place HEAD has been, newest first
  • HEAD@{1}: commit: secondone move ago you were on commit 1992960, "second": the one you "lost"
The Mental Model

A change moves through four places

working diryour edits stagingready to save local repoyour history remoteshared git add git commit git push
Habits That Pay Off

Small commits, clear messages, ignored secrets

You Just Committed Your API Key

Adding it to .gitignore does nothing

This is the single most common security mistake in a student repo.

.gitignore Only Ignores Untracked Files

Once committed, a file stays tracked

This surprises everyone once. .gitignore tells git what to start tracking — it has no effect on something already in the index.

Run in a real repo, not from memory

Commit .env, then add it to .gitignore, then run git ls-files. It is still listed.

echo "SECRET=abc123" > .env git add .env && git commit -m "oops" echo ".env" > .gitignore git add .gitignore && git commit -m "ignore it" git ls-files .env <- STILL TRACKED .gitignore
  • echo "SECRET=abc123" > .envcreate a file .env containing that line (> writes a command's output into a file)
  • &&run the second command only if the first one worked
  • git ls-fileslist every file Git is tracking (watching for changes)
  • .env <- STILL TRACKED.gitignore came too late: the file was already in the index (the staging area)
Untracking Is Not Erasing

The secret is still in every clone

git rm --cached stops future commits from including it. The versions already committed are untouched, and anyone who cloned the repo has them.

Measured in the scratch repo

After untracking, git log -p -- .env still printed the secret twice. History is append-only until you rewrite it.

git rm --cached .env git commit -m "untrack" git ls-files .gitignore # good, not tracked now git log --all -p -- .env | grep -c SECRET 2 # STILL IN HISTORY
  • git rm --cached .envstop tracking the file; --cached keeps it on your disk
  • git log --all -p -- .envevery past version of .env, with its contents (-p)
  • | grep -c SECRET| feeds that output to grep, which counts the lines containing SECRET
  • 2the secret still appears twice in the saved history
What To Actually Do

In this order, and the first step is not git

Ignore Data, Too

Not just secrets

A committed dataset bloats every clone forever — and if it contains personal data, it is in the history under the same rules as a leaked key.

The week-11 connection

DS 227 (the companion course) just covered why deletion is hard. Version control is one of the places a “deleted” file lives on: every old commit still holds it.

# a starting .gitignore .env .env.local data/ *.csv *.parquet __pycache__/ .venv/ .ipynb_checkpoints/ .DS_Store # commit .env.template instead — # same keys, empty values
  • data/a trailing / means a whole folder
  • *.csv* is a wildcard: any file name ending in .csv
  • __pycache__/, .ipynb_checkpoints/folders Python and Jupyter create by themselves
  • .env.templatethe same setting names with blank values: safe to share, shows others what to fill in
Branches

Try something without risking what works

A branch is a movable label on a commit. Making one costs nothing and lets you abandon an experiment without unpicking it. To merge is to bring a branch's commits into another branch.

A branch is a draft copy to try an idea on. If it works, merging pours the good changes back into the main version; if not, you delete the draft.

Branch per piece of work

Not per day, not per person. One branch for "add the regional breakdown", merged when it works.

git switch -c feature/regional-breakdown # ...commit as usual... git switch main git merge feature/regional-breakdown # or throw it away, cost-free git switch main git branch -D feature/regional-breakdown git branch # list them git switch - # previous branch
  • git switch -c feature/…create a new branch (-c) and move onto it
  • git switch maingo back to main, the main line of work
  • git merge feature/…bring that branch's commits into the branch you are on
  • git branch -D feature/…delete the branch, and any work only it had
Commit Messages Are For Your Future Self

They are the only documentation that is never stale

✗ Useless✓ Useful
updatefix region totals double-counted after merge
fix bugpin pandas to 3.0.5 — 2.x changes copy semantics
changesadd duration column + assert it is non-negative
asdfdrop rows with no budget (cannot answer spend question)
Part B

Environments

Pin what you install; run the same anywhere.

You can now keep your code's history. This part makes sure the tools your code needs (pandas, NumPy, …) are the same versions on every computer that runs it.

The Motivation

Different versions, different results

Your code used pandas 2.2; a classmate has 1.5, and it breaks. A virtual environment gives each project its own isolated set of packages, at known versions.

New words

package
an installable library, like pandas (week 5: import pandas)
dependency
a package your code needs in order to run
package manager
the installer: pip (comes with Python) or conda (comes with Anaconda, and can install Python itself)

Isolated

One project's packages can't clash with another's.

Recorded

Exact versions written down, so anyone rebuilds the same setup.

An environment is a separate toolbox for each project. Swapping the hammer in one box does not change the hammer in another.

Create And Activate

A fresh environment per project

Make a .venv folder, activate it, then install. From now on, every pip install lands inside this project only — nothing global is touched.

python -m venv .venv source .venv/bin/activate pip install pandas matplotlib
  • python -m venv .venvcreate a new, empty environment in a folder called .venv (-m venv = run Python's built-in venv tool)
  • source .venv/bin/activateactivate: make this terminal use that environment (Windows: .venv\Scripts\activate)
  • pip install pandas matplotlibdownload and install both packages into .venv only

The conda way (Anaconda users)

conda create -n ds208 python=3.12 pandas then conda activate ds208. Same idea: one named environment per project.

Pin, Do Not Just List

requirements.txt without versions is not reproducible

"pandas" means whatever pandas exists the day someone installs. Between pandas 2 and 3, Copy-on-Write (the week 6 copy rule) became permanent — the same code, different behaviour. A version pin fixes the exact version.

You have already been bitten by this

Week 6 taught the pandas 3 copy rule because the version decides the answer. An unpinned requirement is a moving target.

# not reproducible pandas numpy matplotlib # reproducible pandas==3.0.5 numpy==2.5.3 matplotlib==3.9.2 pip freeze > requirements.txt # or, better, a lock file: uv pip compile requirements.in -o requirements.txt
  • pandas==3.0.5== pins exactly version 3.0.5
  • pip freeze > requirements.txtlist every installed package with its exact version, and write the list into requirements.txt
  • uv pip compile …uv is a newer, faster package manager: it turns your short wish-list (requirements.in) into a lock file that pins every package, including the ones your packages need
One Environment Per Project

So two projects cannot break each other

Installed globally, a package upgrade for one project silently changes another. A per-project environment makes that impossible.

Never commit the environment

.venv/ goes in .gitignore. You commit the recipe (requirements), not the result — it is large and platform-specific.

uv venv # create source .venv/bin/activate # mac/linux .venv\Scripts\activate # windows uv pip install -r requirements.txt # confirm you are where you think which python /path/to/project/.venv/bin/python # wrong answer = you are installing # into the system python
  • uv venvthe same as python -m venv .venv, only faster
  • uv pip install -r requirements.txtinstall exactly what the file lists (-r = read the list from this file)
  • which pythonwhich Python this terminal will run (Windows: where python). It should be inside your project's .venv
The Reproducibility Checklist

What "done" means for the final project

CheckWhy
Runs after Restart & Run AllProves it does not depend on state from deleted cells
Pinned requirements.txtpandas 2 vs 3 changes behaviour, not just speed
Relative paths onlyAn absolute path is the commonest reason a classmate cannot run it
No data or secrets committedBoth live in history forever once pushed
README: what, how to run, where the data comes fromThree short sections; write it on day one
Asserts on the invariants that matterRow counts and impossible values — a failing run beats a wrong number
"It Works On My Machine"

Four causes, in the order to check them

CauseHow to spot itFix
Different package versionpip freeze | grep pandas on both machinesPin it; share the lock file
Different Python versionpython --versionState it in the README; pin in the project config
A file only you haveThey get FileNotFoundErrorRelative paths via pathlib; document how to get the data
Hidden notebook stateWorks for you, fails on a fresh kernelRestart & Run All before you share. Always.
The Recipe File

requirements.txt pins the versions

freeze writes every package and its exact version to a file. Commit it, and anyone can rebuild your environment with one command.

pip freeze > requirements.txt # on any other machine: pip install -r requirements.txt
  • pip freeze > requirements.txtwrite the exact installed versions into the requirements file; commit that file
  • pip install -r requirements.txton another computer, inside a fresh environment: install exactly that list
  • conda env export > environment.ymlthe conda version of the same recipe; rebuild it with conda env create -f environment.yml
Quick Check

Tap to reveal

Which command saves a snapshot of your staged changes to your local history?

A · git add
B · git commit
C · git push
D · git pull

B — git commit.

add stages what to save, commit saves the snapshot locally, and push shares it to the remote. Three distinct steps.

Break

  Five minutes

Back to fix things calmly when they break.

Part C

Debugging

Systematic beats frantic, every time.

You have met error messages all course. This part turns fixing them into a routine: read the error, reproduce it, narrow it down, and check your assumptions.

Don't Panic — Read

The traceback tells you where and what

The last line names the error type and message; the lines above show the path to it. Read it bottom-up — the answer is usually right there, in plain English.

A traceback is the report Python prints when a program stops with an error (an exception); a bug is the mistake in the code that caused it.

File "analysis.py", line 7 df["popn"].mean() KeyError: 'popn' # wrong column name — line 7
  • KeyError: 'popn'read this first. Error type KeyError = you asked for a key (here a column name) that does not exist; the message names it: 'popn'
  • File "analysis.py", line 7where: your file, line 7. Open it at that line
  • df["popn"].mean()the exact line that failed. Check first: is the name spelled as in df.columns?

A full traceback starts with Traceback (most recent call last): and can list many lines inside pandas' own files. Skip those: look for the lines that name your file.

See What's Really There

Print the values; pause to inspect

Most bugs are a value that isn't what you assumed. A quick print checks it; breakpoint() pauses execution so you can poke around live.

print(f"shape: {df.shape}") print(df.columns.tolist()) breakpoint() # drops into the debugger
  • print(f"shape: {df.shape}")an f-string fills in the value inside { }: for the flood data it prints shape: (34079, 15), i.e. rows, columns
  • df.columns.tolist()the real column names as a plain list: compare the spelling with what you typed
  • breakpoint()pause here and open the debugger, a tool for looking inside a paused program
A Method, Not A Panic

Reproduce, isolate, check

breakpoint()

Stop time and look around

A print tells you one value you thought to ask for. A breakpoint lets you ask anything, at that moment, including questions you only think of once you are there.

Built in since Python 3.7

No import needed. Type breakpoint() on the line before the failure and run normally.

A breakpoint is a pause button on a film: the frame freezes and you can look at everything in it. Stepping is advancing one frame at a time.

def clean(df): breakpoint() # execution stops here return df.dropna() # at the (Pdb) prompt: p df.shape print an expression n next line s step into a call c continue l list code around here q quit
  • (Pdb)the debugger's prompt: type a one-letter command, press Enter
  • p df.shapeprint any value, as it is right now
  • n vs srun the next line; or step into the function that line calls
  • c, qcarry on to the next breakpoint; or stop the program
Binary Search Your Bug

Halve the suspect region until it is one line

When a 60-line pipeline produces a wrong answer, do not read all 60. Check the middle, then discard half.

Six checks for sixty lines

Each check halves what is left. This is the same trick as guessing a number from 1 to 100 by always asking “higher or lower?” (called binary search), applied to your own code.

# where did the row count go wrong? df = load() ; print(len(df)) 34079 df = clean(df) ; print(len(df)) 34079 df = derive(df) ; print(len(df)) 34079 df = merge_regions(df); print(len(df)) 68158 <- # found it. a duplicated join key. # you never read merge_regions until # the counter pointed at it
  • … ; print(len(df)); puts two statements on one line: do the step, then print the row count
  • 34079after load, clean and derive: still one row per project, as expected
  • 68158 <-exactly double: the merge matched every row twice, because the region table listed each region twice (a duplicated join key, the column that matches rows, week 7)
assert, Where You Assume

Fail at the cause, not three steps later

The expensive bugs are the ones that stay silent and surface far from their origin. An assertion (assert) converts a silent wrong answer into a loud stop.

Assert the invariant, not the value

"the row count did not change" and "durations are not negative" — the properties that must hold if the step worked.

before = len(df) df = df.merge(regions, on="region", how="left") assert len(df) == before, f"row count changed: {before} -> {len(df)}" assert df.duration.min() >= 0, "negative duration" AssertionError: negative duration # the flood data really has one: a project # "completed" 17 days before it started. # an assert finds it every run.
  • assert condition, messageif the condition is True, nothing happens; if False, stop with AssertionError and show the message
  • len(df) == beforethe invariant: a left merge should keep the row count the same
  • df.duration.min() >= 0duration = days from start to completion (the column built in week 8); the smallest must not be below 0
  • AssertionError: negative durationreading it: AssertionError = one of your own checks failed; the message says which one
logging Beats print

Because you can turn it down without deleting it

Debug prints get deleted, then re-added next time the bug appears. Logging lets the same lines stay and go quiet.

Levels are the point

Set INFO normally, drop to DEBUG when something breaks. No edits, no re-adding prints.

import logging logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s") log = logging.getLogger(__name__) log.debug("rows after clean: %s", len(df)) log.info("loaded %s projects", len(df)) log.warning("%s rows had no budget", n) # each line starts with its level, # so you can search for WARNING
  • logging.basicConfig(level=…INFO, …)show messages of level INFO and above, formatted as "LEVEL message"
  • log = logging.getLogger(__name__)make a logger named after this file
  • log.debug(…) / log.info(…)at level INFO the debug line is hidden; the info line prints INFO loaded 34079 projects (%s is filled with len(df))
  • log.warning(…)with n = 12 it prints WARNING 12 rows had no budget. Levels run DEBUG < INFO < WARNING < ERROR; lower the level to see more
The Minimal Reproducible Example

Cutting it down usually solves it

A minimal reproducible example (MRE) is the shortest piece of code, with the smallest data, that still shows the bug.

Your Turn · 8 min

Break it, then recover

1 · Commit a "secret"

In a terminal, make a repo (mkdir demo, cd demo, git init), commit a file with a fake key. Add it to .gitignore. Confirm with git ls-files that it is still tracked.

2 · Untrack and check history

git rm --cached, commit, then git log -p. Find the fake key still there.

3 · Lose a commit and get it back

reset --hard HEAD~1, then recover it from git reflog.

4 · Assert your pipeline

Add a row-count assert around a merge. Make it fail deliberately, and read where it stops.

Quick Check

Tap to reveal

You committed .env with a live API key, then added .env to .gitignore and pushed. What is your situation?

A · Fixed — .gitignore stops it being tracked
B · The key is still tracked and still in history; rotate it first, then clean the history
C · Fixed once you run git rm --cached .env
D · Only a problem if the repository is public
B. .gitignore has no effect on an already-tracked file — git ls-files still lists it. git rm --cached (C) stops future commits but leaves every past version in history. Rotating the key is the only step that definitely works, and it comes first; history rewriting is slow and needs coordination. D is wrong: private repos get cloned, forked and backed up too.
Quick Check

Tap to reveal

Your notebook runs perfectly for you and fails on a classmate’s machine at cell 3. What do you check first?

A · Their internet connection
B · Whether it still works for you after Restart & Run All
C · Reinstall their Python
D · Send them your .venv folder
B. A notebook keeps state from cells you have since edited or deleted, so it can depend on code that no longer exists in the file. Restart & Run All is the first check because it is the most common cause and takes ten seconds. D is never right — environments are large and platform-specific; share the pinned requirements instead.
Putting It Together

A KeyError, found in three lines

The error named a missing column. Printing the real column names showed the typo. Read, check, fix — no guessing, no flailing.

df["popn"].mean() KeyError: 'popn' df.columns # ['city','region','pop'] df["pop"].mean() # fixed
  • KeyError: 'popn'last line first: no column called popn
  • df.columnscheck the assumption: the real names are city, region, pop
  • df["pop"]use the real name; the error is gone
Quick Check

Tap to reveal

When reading a traceback, which line usually names the actual error?

A · the very first line
B · the last line
C · the middle line
D · it's random

B — the last line.

The bottom line gives the error type and message; the lines above trace how you got there. Read bottom-up: what, then where.

This Week's Lab

Version it, pin it, debug it

In the browser lab you'll read tracebacks and fix planted bugs, name error types, guard code with assert, and record the exact package versions your code ran with. The Git and environment steps run in your own terminal (the Your Turn activity). ~45 minutes.

You'll practise

Reading tracebacks, try/except (week 4), assert, and __version__ (e.g. pd.__version__ gives '2.2.0', the installed version); plus git add/commit/push, venv and pip freeze in a terminal.

Stretch, if you want

Write a .gitignore that excludes data/ and .venv/.

Recap

Track, pin, debug

Glossary · 1 of 2

Git words, one line each

Words for today

reproducibility
anyone can rerun your work and get the same result
version control
saving named snapshots of files to return to or share (Git)
repository (repo)
a folder Git tracks, plus its whole history
commit
one saved snapshot, with a message saying why
branch
a separate line of commits for trying something safely
merge
bring one branch's commits into another
remote
the shared copy of the repo, e.g. on GitHub
push / pull
upload your commits / download other people's

Also new today

terminal
text window where you type commands such as git and pip
stage
put changes in the pile for the next commit (git add)
working directory
your project folder as you see it now
SHA
a commit's ID, like 6b80d5e
diff
the line-by-line changes between two versions
HEAD
the commit you are on now; HEAD~1 is the one before
revert / reset
undo by adding a new commit / by moving back and discarding
reflog
local diary of every commit HEAD has pointed to
.gitignore
list of files Git should never start tracking
clone
a full copy of a repo, history included
Glossary · 2 of 2

Environment and debugging words, one line each

Words for today

dependency
a package your code needs in order to run
environment
Python plus the exact installed packages; virtual = one per project
package manager (pip / conda)
the tool that installs packages
requirements file
requirements.txt: the packages a project needs, one per line
version pin
the exact version to install: pandas==3.0.5
debugger
tool that pauses a program so you can inspect it (pdb)
breakpoint
the line where the program pauses: breakpoint()
stepping
running a paused program one line at a time (n, s)
assertion
assert: stop with an error if a condition is False

Also new today

package
an installable library, such as pandas
lock file
pins every package, including the packages your packages need
activate
make the terminal use a given environment
traceback
Python's error report: read the last line first, then the line number
bug
the mistake in the code that causes a wrong result or error
invariant
a fact that must always hold, e.g. row count unchanged
logging
messages with levels you can turn up or down
kernel
the running Python behind a notebook
minimal reproducible example
shortest code and data that still shows the bug
Before Next Week

Practice & the project

Next Week

Integrative Project & Presentations

One dataset, the whole toolkit: load it, clean it, analyse it, visualise it — then present the story to the class.

DS 208 · Programming for Data Science