Code someone else can run next year and get the same number — with Git for history, environments for dependencies, and calm debugging.
Programming for Data Science · University of the Philippines Cebu
Version history: undo anything, and work with others safely.
Pin your dependencies so the code runs the same anywhere.
A calm, systematic way to find and fix what broke.
The course's motto lives here: code someone else can run next year and get the same number. These are the habits that make that true.
Track a change in Git and lock your project's dependencies.
Someone else can rerun your work, on their computer, and get the same result.
e.g. a classmate runs your notebook next year and sees the same total
A system that saves named snapshots of your files so you can go back to, compare or share any of them. Git is the one we use.
e.g. "show me the code as it was last Tuesday"
A project folder that Git is tracking, together with its whole saved history.
e.g. your ds208-project/ folder after git init
One saved snapshot of the project, with a message saying what changed and why.
e.g. git commit -m "Fix region totals"
A separate line of commits where you can try something without touching the working version.
e.g. a branch called feature/regional-breakdown
Bring the commits from one branch into another, combining the two lines of work.
e.g. git merge feature/regional-breakdown while on main
A copy of the repository on another computer, usually a website like GitHub, that everyone shares.
e.g. github.com/you/ds208-project
Push sends your new commits up to the remote; pull brings other people's commits down to you.
e.g. git push after committing; git pull before starting
A package your code needs in order to run.
e.g. your analysis depends on pandas and matplotlib
The Python plus the exact set of installed packages that your code runs with. A virtual environment is a private one per project.
e.g. a .venv folder inside your project
The tool that downloads and installs packages. pip comes with Python; conda comes with Anaconda.
e.g. pip install pandas
A text file, requirements.txt, listing every package the project needs, one per line.
e.g. a line reading pandas==3.0.5
Writing down the exact version to install, so everyone gets the same one.
e.g. ==3.0.5 in pandas==3.0.5
A tool that pauses your running program so you can look at every variable and move forward line by line.
e.g. Python's built-in pdb
A marked line where the program pauses and hands control to the debugger.
e.g. writing breakpoint() on the line before the crash
Running a paused program one line at a time, watching what each line does.
e.g. typing n (next line) at the (Pdb) prompt
A line that checks something must be true and stops the program with an error if it is not.
e.g. assert len(df) > 0, "no rows loaded"
You can load, analyse, and plot (weeks 5–10). But if only you can run it, on only your laptop, it isn't science yet. Reproducibility (anyone can rerun it and get the same result) is what makes work trustworthy.
A good recipe lists exact amounts and steps, so anyone can bake the same cake. Today is about writing your analysis like that recipe, not like "a bit of this, a pinch of that".
Lost history, mismatched versions, bugs you can't retrace.
Track changes, pin dependencies, debug methodically.
A save-point machine for your code.
You already save files. This part adds version control: saving named snapshots you can return to, compare, and share with others.
Git records snapshots of your project over time. You can undo, see who changed what and why, and merge work from several people without chaos.
A folder Git tracks is a repository (repo); each saved snapshot is a commit.
Into a terminal (the command line): a text window where you type instructions to your computer. Terminal on macOS/Linux, PowerShell or Git Bash on Windows. Not into Python or a notebook cell. Run them from inside your project folder.
Roll back to any past snapshot in seconds.
Many people, one project, no overwriting each other.
Git is like the version history in Google Docs, except you decide when a version is saved, and you write a note on each one saying why.
Stage the files you changed, save a snapshot with a message, and send it to the shared copy. Three commands you'll run dozens of times a day.
To stage a file is to put its changes in the "next snapshot" pile; to push is to upload your commits to the remote, the shared copy (e.g. on GitHub).
git add analysis.pystage: include this file's current changes in the next commitgit commit -m "…"save everything staged as one commit; -m attaches the message in quotesgit pushupload your new commits to the remote so others can get themVersion control earns its keep when something breaks and you need to know what changed. These are how you ask.
The default format shows four lines per commit. You want one, so you can scan fifty at once.
git log --oneline -10the last 10 commits, one line each: short ID, then message6b80d5ea commit's ID (its SHA, a unique fingerprint); type it to point at that commit-- path/to/fileonly the commits that touched that one filegit show 8e2216ewhat that one commit changedgit diffa diff is the line-by-line changes: here, edits you have not staged yet; --staged shows what the next commit will savegit blame file.pyfor each line, which commit (and person) last changed it| Command | What it undoes | History | Safe to use after pushing? |
|---|---|---|---|
git restore <file> | Uncommitted edits to that file | Untouched | Yes — nothing was committed |
git revert <sha> | A commit, by adding an inverse commit | Preserved, moves forward | Yes — this is the shared-branch answer |
git reset --hard <sha> | Everything back to that commit | Rewritten | No — it breaks everyone else’s clone |
The rule: revert in public, reset in private. Once a commit is pushed and someone has pulled it, rewriting it creates work for them; reverting it does not.
<file>, <sha>A hard reset looks like the commit is gone — it is not in git log any more. But git kept the pointer for weeks.
Run in a scratch repo: after reset --hard HEAD~1 the discarded commit is the second line of git reflog.
HEAD~1HEAD is the commit you are on now; ~1 means "one before it"git reset --hard …move back to that commit and throw away everything after itgit reflogGit's private diary of every place HEAD has been, newest firstHEAD@{1}: commit: secondone move ago you were on commit 1992960, "second": the one you "lost"Each command moves your change one step right. add picks what to
save, commit saves it, push shares it. The working
directory is your project folder as you see it; the staging area is
the pile for the next commit.
Like sending a parcel: put items in the box (add), seal and label it
(commit), hand it to the courier (push). pull is receiving the parcels others
sent.
git pull brings others' changes back down before you start.
One idea per commit. Easy to review, easy to undo.
"Fix off-by-one in slice," not "stuff." Future you is the reader.
.gitignoreNever commit data, secrets, or .venv. List them to ignore them.
A committed API key or password is a leak — even after you delete it, it lives
on in the history. Ignore secrets from day one. .gitignore is a plain text file in
the repo listing files and folders Git should never start tracking; a secret
is anything that grants access, such as a password or API key (week 9).
Keys go in environment variables and a .gitignore, never in a commit.
This is the single most common security mistake in a student repo.
This surprises everyone once. .gitignore tells git what to start tracking — it has no effect on something already in the index.
Commit .env, then add it to .gitignore, then run git ls-files. It is still listed.
echo "SECRET=abc123" > .envcreate a file .env containing that line (> writes a command's output into a file)&&run the second command only if the first one workedgit ls-fileslist every file Git is tracking (watching for changes).env <- STILL TRACKED.gitignore came too late: the file was already in the index (the staging area)git rm --cached stops future commits from including it. The versions already committed are untouched, and anyone who cloned the repo has them.
After untracking, git log -p -- .env still printed the secret twice. History is append-only until you rewrite it.
git rm --cached .envstop tracking the file; --cached keeps it on your diskgit log --all -p -- .envevery past version of .env, with its contents (-p)| grep -c SECRET| feeds that output to grep, which counts the lines containing SECRET2the secret still appears twice in the saved historyAssume it is compromised the moment it was pushed. Revoke the key and issue a new one — this is the only step that definitely works.
git filter-repo (or BFG) rewrites every commit. It changes every SHA, so coordinate with anyone who has cloned.
.env in .gitignore from the first commit, a committed .env.template with blank values, and a secret scanner in CI.
Step 1 first, always. Rewriting history is slow and coordinated; rotation is immediate and complete. A cleaned repo with a live leaked key is still compromised.
1992960A committed dataset bloats every clone forever — and if it contains personal data, it is in the history under the same rules as a leaked key.
DS 227 (the companion course) just covered why deletion is hard. Version control is one of the places a “deleted” file lives on: every old commit still holds it.
data/a trailing / means a whole folder*.csv* is a wildcard: any file name ending in .csv__pycache__/, .ipynb_checkpoints/folders Python and Jupyter create by themselves.env.templatethe same setting names with blank values: safe to share, shows others what to fill inA branch is a movable label on a commit. Making one costs nothing and lets you abandon an experiment without unpicking it. To merge is to bring a branch's commits into another branch.
A branch is a draft copy to try an idea on. If it works, merging pours the good changes back into the main version; if not, you delete the draft.
Not per day, not per person. One branch for "add the regional breakdown", merged when it works.
git switch -c feature/…create a new branch (-c) and move onto itgit switch maingo back to main, the main line of workgit merge feature/…bring that branch's commits into the branch you are ongit branch -D feature/…delete the branch, and any work only it had| ✗ Useless | ✓ Useful |
|---|---|
update | fix region totals double-counted after merge |
fix bug | pin pandas to 3.0.5 — 2.x changes copy semantics |
changes | add duration column + assert it is non-negative |
asdf | drop rows with no budget (cannot answer spend question) |
The test: in six months, scanning git log --oneline, could you find the commit that introduced a behaviour? Say what changed and why — the diff already shows how.
Pin what you install; run the same anywhere.
You can now keep your code's history. This part makes sure the tools your code needs (pandas, NumPy, …) are the same versions on every computer that runs it.
Your code used pandas 2.2; a classmate has 1.5, and it breaks. A virtual environment gives each project its own isolated set of packages, at known versions.
import pandas)pip (comes with Python) or conda (comes with Anaconda, and can install Python itself)One project's packages can't clash with another's.
Exact versions written down, so anyone rebuilds the same setup.
An environment is a separate toolbox for each project. Swapping the hammer in one box does not change the hammer in another.
Make a .venv folder, activate it, then install. From now on, every
pip install lands inside this project only — nothing global is touched.
python -m venv .venvcreate a new, empty environment in a folder called .venv (-m venv = run Python's built-in venv tool)source .venv/bin/activateactivate: make this terminal use that environment (Windows: .venv\Scripts\activate)pip install pandas matplotlibdownload and install both packages into .venv onlyconda create -n ds208 python=3.12 pandas then conda activate ds208. Same idea: one named environment per project.
"pandas" means whatever pandas exists the day someone installs. Between pandas 2 and 3, Copy-on-Write (the week 6 copy rule) became permanent — the same code, different behaviour. A version pin fixes the exact version.
Week 6 taught the pandas 3 copy rule because the version decides the answer. An unpinned requirement is a moving target.
pandas==3.0.5== pins exactly version 3.0.5pip freeze > requirements.txtlist every installed package with its exact version, and write the list into requirements.txtuv pip compile …uv is a newer, faster package manager: it turns your short wish-list (requirements.in) into a lock file that pins every package, including the ones your packages needInstalled globally, a package upgrade for one project silently changes another. A per-project environment makes that impossible.
.venv/ goes in .gitignore. You commit the recipe (requirements), not the result — it is large and platform-specific.
uv venvthe same as python -m venv .venv, only fasteruv pip install -r requirements.txtinstall exactly what the file lists (-r = read the list from this file)which pythonwhich Python this terminal will run (Windows: where python). It should be inside your project's .venv| Check | Why |
|---|---|
| Runs after Restart & Run All | Proves it does not depend on state from deleted cells |
Pinned requirements.txt | pandas 2 vs 3 changes behaviour, not just speed |
| Relative paths only | An absolute path is the commonest reason a classmate cannot run it |
| No data or secrets committed | Both live in history forever once pushed |
| README: what, how to run, where the data comes from | Three short sections; write it on day one |
| Asserts on the invariants that matter | Row counts and impossible values — a failing run beats a wrong number |
Every item is something that has already bitten this course in a previous week. Run the checklist on a fresh clone in a fresh environment — that is the only test that actually proves it.
data/flood.csv, measured from the project folder; not C:\Users\ana\…| Cause | How to spot it | Fix |
|---|---|---|
| Different package version | pip freeze | grep pandas on both machines | Pin it; share the lock file |
| Different Python version | python --version | State it in the README; pin in the project config |
| A file only you have | They get FileNotFoundError | Relative paths via pathlib; document how to get the data |
| Hidden notebook state | Works for you, fails on a fresh kernel | Restart & Run All before you share. Always. |
The last one is the most common and the least suspected. A notebook accumulates variables from cells you have since edited or deleted — it can depend on code that no longer exists anywhere.
pip freeze | grep pandas| feeds the package list to grep, which keeps only lines containing "pandas"pathlibrequirements.txt pins the versionsfreeze writes every package and its exact version to a file. Commit
it, and anyone can rebuild your environment with one command.
pip freeze > requirements.txtwrite the exact installed versions into the requirements file; commit that filepip install -r requirements.txton another computer, inside a fresh environment: install exactly that listconda env export > environment.ymlthe conda version of the same recipe; rebuild it with conda env create -f environment.ymlWhich command saves a snapshot of your staged changes to your local history?
git addgit commitgit pushgit pullB — git commit.
add stages what to save,
commit saves the snapshot locally, and push shares it to the
remote. Three distinct steps.
add → stage, commit → save, push → share. Keep the three straight.
commit writes to history; push shares it.
Back to fix things calmly when they break.
Systematic beats frantic, every time.
You have met error messages all course. This part turns fixing them into a routine: read the error, reproduce it, narrow it down, and check your assumptions.
The last line names the error type and message; the lines above show the path to it. Read it bottom-up — the answer is usually right there, in plain English.
A traceback is the report Python prints when a program stops with an error (an exception); a bug is the mistake in the code that caused it.
KeyError: 'popn'read this first. Error type KeyError = you asked for a key (here a column name) that does not exist; the message names it: 'popn'File "analysis.py", line 7where: your file, line 7. Open it at that linedf["popn"].mean()the exact line that failed. Check first: is the name spelled as in df.columns?A full traceback starts with
Traceback (most recent call last): and can list many lines inside pandas' own
files. Skip those: look for the lines that name your file.
Most bugs are a value that isn't what you assumed. A quick print
checks it; breakpoint() pauses execution so you can poke around live.
print(f"shape: {df.shape}")an f-string fills in the value inside { }: for the flood data it prints shape: (34079, 15), i.e. rows, columnsdf.columns.tolist()the real column names as a plain list: compare the spelling with what you typedbreakpoint()pause here and open the debugger, a tool for looking inside a paused programMake the error happen again, on purpose. Find the smallest input that triggers it. A reliable bug is a solvable bug.
Comment out half. Narrow down to the exact line at fault.
Print the value you're sure about. It's often the one that's wrong.
Stuck? Explain the problem aloud, line by line — the "rubber duck." Saying it often reveals the flaw before anyone answers.
Paste the exact error message into a search — someone has hit it before.
A print tells you one value you thought to ask for. A breakpoint lets you ask anything, at that moment, including questions you only think of once you are there.
No import needed. Type breakpoint() on the line before the failure and run normally.
A breakpoint is a pause button on a film: the frame freezes and you can look at everything in it. Stepping is advancing one frame at a time.
(Pdb)the debugger's prompt: type a one-letter command, press Enterp df.shapeprint any value, as it is right nown vs srun the next line; or step into the function that line callsc, qcarry on to the next breakpoint; or stop the programWhen a 60-line pipeline produces a wrong answer, do not read all 60. Check the middle, then discard half.
Each check halves what is left. This is the same trick as guessing a number from 1 to 100 by always asking “higher or lower?” (called binary search), applied to your own code.
… ; print(len(df)); puts two statements on one line: do the step, then print the row count34079after load, clean and derive: still one row per project, as expected68158 <-exactly double: the merge matched every row twice, because the region table listed each region twice (a duplicated join key, the column that matches rows, week 7)The expensive bugs are the ones that stay silent and surface far from their origin. An assertion (assert) converts a silent wrong answer into a loud stop.
"the row count did not change" and "durations are not negative" — the properties that must hold if the step worked.
assert condition, messageif the condition is True, nothing happens; if False, stop with AssertionError and show the messagelen(df) == beforethe invariant: a left merge should keep the row count the samedf.duration.min() >= 0duration = days from start to completion (the column built in week 8); the smallest must not be below 0AssertionError: negative durationreading it: AssertionError = one of your own checks failed; the message says which oneDebug prints get deleted, then re-added next time the bug appears. Logging lets the same lines stay and go quiet.
Set INFO normally, drop to DEBUG when something breaks. No edits, no re-adding prints.
logging.basicConfig(level=…INFO, …)show messages of level INFO and above, formatted as "LEVEL message"log = logging.getLogger(__name__)make a logger named after this filelog.debug(…) / log.info(…)at level INFO the debug line is hidden; the info line prints INFO loaded 34079 projects (%s is filled with len(df))log.warning(…)with n = 12 it prints WARNING 12 rows had no budget. Levels run DEBUG < INFO < WARNING < ERROR; lower the level to see moreA minimal reproducible example (MRE) is the shortest piece of code, with the smallest data, that still shows the bug.
Delete everything the error does not need. Most bugs become obvious at about line five.
Three hand-written rows instead of 34,079. If it still fails, the data was never the problem.
You now have something someone can run in ten seconds — which is the difference between a question that gets answered and one that does not.
The useful part: you will usually find the bug while cutting it down. Every piece you remove is a hypothesis tested. The MRE is debugging, and being able to ask a good question is the by-product.
In a terminal, make a repo (mkdir demo, cd demo, git init), commit a file with a fake key. Add it to .gitignore. Confirm with git ls-files that it is still tracked.
git rm --cached, commit, then git log -p. Find the fake key still there.
reset --hard HEAD~1, then recover it from git reflog.
Add a row-count assert around a merge. Make it fail deliberately, and read where it stops.
Step 1 is the one to actually run. Reading that .gitignore does not untrack a committed file is not the same as watching it happen to you.
You committed .env with a live API key, then added .env to .gitignore and pushed. What is your situation?
git rm --cached .env.gitignore has no effect on an already-tracked file — git ls-files still lists it. git rm --cached (C) stops future commits but leaves every past version in history. Rotating the key is the only step that definitely works, and it comes first; history rewriting is slow and needs coordination. D is wrong: private repos get cloned, forked and backed up too.Your notebook runs perfectly for you and fails on a classmate’s machine at cell 3. What do you check first?
.venv folderThe error named a missing column. Printing the real column names showed the typo. Read, check, fix — no guessing, no flailing.
KeyError: 'popn'last line first: no column called popndf.columnscheck the assumption: the real names are city, region, popdf["pop"]use the real name; the error is goneWhen reading a traceback, which line usually names the actual error?
B — the last line.
The bottom line gives the error type and message; the lines above trace how you got there. Read bottom-up: what, then where.
Read tracebacks bottom-up. The last line is the headline.
Error type + message = last line.
In the browser lab you'll read tracebacks and fix planted bugs, name error types,
guard code with assert, and record the exact package versions your code ran with.
The Git and environment steps run in your own terminal (the Your Turn activity). ~45 minutes.
Reading tracebacks, try/except (week 4), assert, and
__version__ (e.g. pd.__version__ gives '2.2.0', the installed version); plus git add/commit/push, venv and
pip freeze in a terminal.
Write a .gitignore that excludes data/ and .venv/.
add → commit → push; small commits, clear messages, ignore secrets.
venv to isolate; requirements.txt to pin versions.
Read the traceback; reproduce, isolate, check assumptions.
One sentence: record your history, pin your dependencies, and debug with a method — so your work runs the same for everyone.
Hand a project to someone else and have it just work.
git and pipgit add)6b80d5eHEAD~1 is the one beforerequirements.txt: the packages a project needs, one per linepandas==3.0.5pdb)breakpoint()n, s)assert: stop with an error if a condition is FalseFinish the Week 11 lab and submit it. Put your integrative project under Git from the first commit.
GitHub's "Git Handbook" — docs.github.com/en/get-started/using-git.
Everything here is linked on the course page beside this deck.
Next week: the integrative project — every skill in this course, on one dataset.
One dataset, the whole toolkit: load it, clean it, analyse it, visualise it — then present the story to the class.
DS 208 · Programming for Data Science