Why code beats clicks, what the Python data stack looks like, and how to get your machine — or the cloud — ready to compute.
3 units · Distance education (asynchronous online)
This deck is designed for asynchronous viewing: dark "pause" slides mark good stopping points between parts. Part C works best with a second screen or phone showing these slides while you type.
≈3 hours of material — matching the course's 3 weekly contact hours. Don't binge it; the pauses are load-bearing.
What code gives you that spreadsheets and point-and-click tools fundamentally can't.
A guided map of NumPy, pandas, matplotlib, and friends — who does what.
Getting a working environment today: local install or zero-install in the cloud.
Goal for the week: run your first notebook cell on real data — before next Monday.
Work at your own pace, but the Week 1 setup checkpoint is firm — everything else builds on it.
You do not need to memorise these now. Each one is explained again on the slide where it first matters, and they are all collected in a glossary near the end.
A file of code — written instructions — that Python runs from the top line to the bottom.
e.g. a file called clean_survey.py
A bundle of ready-made code someone else wrote, which you load into your own work.
e.g. pandas (tables), NumPy (numbers), matplotlib (charts)
The line that loads a library so your code can use it.
e.g. import pandas as pd — load pandas, call it pd for short
A document of boxes that mix text and runnable code. Jupyter and Colab show notebooks.
e.g. week01_hello_data.ipynb
One box in a notebook. Shift+Enter runs it; the result appears underneath.
e.g. a cell containing 1 + 1 shows 2
The Python program running behind a notebook. It remembers everything you have run so far.
e.g. “Restart kernel” wipes that memory clean
A name that holds a value (a piece of data), like a labelled box.
e.g. price = 100 puts 100 in a box labelled price
A named piece of work. You call it by writing its name and brackets; inputs go inside the brackets.
e.g. print("hi") calls the function print
Python stopping because it cannot do what a line asks. Normal, expected, and fixable.
e.g. a misspelt name gives a NameError
The report printed with an error. Read the last line first: it names the problem.
e.g. NameError: name 'pirce' is not defined
In October 2020, Public Health England lost track of nearly 16,000 positive COVID-19 cases — results vanished because an old Excel format silently truncated rows beyond its limit.
Not "spreadsheets are bad" — they're great for small tables. The lesson is that invisible, unrepeatable data handling fails silently at scale.
A month later, nobody — including you — can say exactly which cells were edited, filtered, or pasted.
New month, new file, same 40 clicks — every repetition is a fresh chance for a new mistake.
Millions of rows, nested JSON, joins across ten files, statistical models — the grid runs out of headroom.
Reinhart & Rogoff's influential 2010 paper "Growth in a Time of Debt" argued growth collapses when public debt passes 90% of GDP — and was cited to justify austerity worldwide. In 2013, graduate student Thomas Herndon tried to replicate it and found, among other issues, an Excel formula that silently omitted five countries.
The error surfaced only because Herndon requested the actual spreadsheet and re-ran the analysis. No script existed to audit — the analysis lived in cell selections.
Three superpowers: reproducibility, scale, automation.
You already know how to analyse data by clicking through a spreadsheet. This part shows what you gain when those clicks are written down as code instead.
Code is written instructions for the computer. A script is a file of code that runs from top to bottom: a complete, ordered, re-runnable record of every step from raw file to final figure. Hand it to a colleague — or your future self — and the result reappears exactly.
Reproducibility is the currency of graduate research. Reviewers and advisers increasingly expect runnable code, not screenshots.
You are not expected to write this yet. Read it like a recipe card: each line is one step, in order, and anyone can follow it again.
# The whole analysis…a comment: a note for humans. Python ignores everything after #import pandas as pdload the pandas library and call it pd for shortdf = pd.read_csv(…)read the CSV file into a table and store it under the name dfdf.dropna(subset=["program"])throw away rows where the program column is blankdf.groupby("program").size()count how many rows there are for each programsummary.to_csv(…)save those counts to a new file. No output on screen — the result is the file| Task | By hand | With a script |
|---|---|---|
| Clean 1 survey file | 20 minutes of careful clicking | Write once: ~15 lines |
| Clean 300 survey files | Weeks, with drifting rules | The same 15 lines in a loop |
| Redo everything after a data fix | Start over, hope for consistency | Re-run: one command |
| Monthly report, every month | A recurring calendar dread | Scheduled job, done before coffee |
The pattern: effort moves from every time to one time. That trade is the entire economic case for this course. (A loop tells the computer to repeat the same lines once per item — here, once per file. A scheduled job is a script the computer starts by itself at a set time.)
The first script takes longer than the first manual pass. The payoff starts at repetition two.
It does exactly what you wrote, never what you meant. Every bug (a mistake that makes code do the wrong thing) is a gap between those two — usually a small, findable one.
Professionals see error messages constantly; they've just stopped reading them as failure and started reading them as directions.
Any analysis, however grand, is small steps composed: load, filter, group, plot. If a step is confusing, split it again.
Expect the first three weeks to feel slow and mechanical. That's the fingers learning; the speed arrives suddenly around Week 5 when pandas clicks.
Tolerance for being briefly, fixably wrong — several times per hour.
Import = get the data into Python. Tidy = arrange it so each row is one thing and each column one measurement. Transform = filter, compute and summarise. Visualize = draw it. Model = fit a statistical model. Communicate = explain what you found to someone else.
Adapted from Wickham & Grolemund's R for Data Science — the workflow is language-agnostic; we'll walk it in Python.
Import/tidy → Weeks 5–7 · visualize → Week 8 · every stage → your Week 12 project.
A colleague proudly shows you their analysis: a beautifully formatted Excel workbook where they manually filtered, copied cleaned rows to a new sheet, and built a chart. The numbers are correct. From a programming-for-data-science standpoint, what's the core problem?
Commit to an answer, then click and hold (or tap) to reveal
C — it's a reproducibility problem, not a correctness problem.
The work may be flawless today, but the manual steps are invisible: they can't be audited (Reinhart–Rogoff), can't be re-run when data updates, and can't scale to the next 300 files. The result isn't wrong — it's unrepeatable.
Good place to stop for today, or stretch for five minutes. Part B is a guided tour of the Python libraries you'll live in all semester.
One language, a division of labor.
You now know why we write code. This part shows what we write it with: one language, Python, plus a handful of libraries that each do one job.
| Language | Strengths | Our take |
|---|---|---|
| Python | Readable, huge ecosystem, general-purpose (web, automation, ML) | Our vehicle — one language covers the whole workflow |
| R | Statistics-first, superb plotting (ggplot2) | Excellent; learn it later if your field favors it |
| SQL | The lingua franca of databases | Not a rival — you'll meet it in Week 9 |
| Julia | Speed for numerical computing | Promising, smaller ecosystem |
A programming language is a strict, written vocabulary for giving a computer instructions; its ecosystem is all the libraries and tools people have built for it. Skills transfer: the concepts you learn — thinking in tables, operating on whole columns at once (vectorized operations), tidy workflows — carry to any of these languages.
That's the design assumption. Weeks 2–4 build the language from the ground up.
NumPy = fast arrays. pandas = labeled tables. 80% of your code will touch them. You never install “data science” — you assemble it from these parts, week by week.
NumPy's array is a row of numbers, all of the same kind, that NumPy can do maths on in one go. It replaces the loop: you write the operation once and it applies to every element (every number in the array), very fast. This idea — vectorization, one operation applied to a whole array — is the single biggest mental shift from ordinary programming.
It is like typing a formula once in Excel and having it fill the whole column — without even dragging it down.
Today's goal is only to recognize the shape of the code. Week 5 makes you fluent.
rain = np.array([…])make an array of six rainfall numbers and name it rainrain * 0.1multiply every number by 0.1 at once — no loop neededrain.mean()the average of all six: about 131.6 mmrain > 100ask “over 100?” of each number; each answer is True (yes) or False (no)rain[rain > 100]keep only the numbers where the answer was Truepandas puts labels on NumPy's speed. Its table is called a DataFrame: rows and named columns, like a spreadsheet you control with code. The questions you'd click through in Excel become one readable line each.
A CSV file (“comma-separated values”) is a plain-text table: one row per line, commas between the columns. Excel can save one; pandas can read one.
Weeks 6–7 are pure pandas. Most working data scientists spend most hours exactly here.
df = pd.read_csv("…csv")read the CSV file into a DataFrame called dfdf.groupby("barangay")["no_show"].mean()for each barangay, the average of the no_show column — a pivot table in one linedf[df["distance_km"] > 4]keep only the rows where distance is over 4 km.sort_values("age")then put those rows in order of agematplotlib draws; seaborn (built on top) makes the common statistical plots pretty by default. Every chart is code — which means every chart is reproducible and revisable.
By Week 8 you'll treat a figure the way you treat a paragraph: drafted, revised, and defended.
import matplotlib.pyplot as pltload matplotlib’s drawing tools, nicknamed pltmonthly = df.groupby("month")…sum()total visits per month, from a table df loaded earliermonthly.plot(kind="bar")draw those totals as a bar chartplt.title(…) / plt.show()add a title, then display the finished chartYou have a CSV of 200,000 e-wallet transactions and want the average amount per province. Which layer of the stack does the heavy lifting?
Commit to an answer, then click and hold (or tap) to reveal
C — this is the pandas shape: labeled table, grouped summary.
A works but is slow to write and easy to get wrong; B rebuilds what pandas already gives you; D confuses drawing with computing. Matching the task to the layer is half of fluency — and it's learnable.
Next part is hands-on: have your laptop ready and ~1 hour of focus. You'll leave it with a working environment — this week's actual deliverable.
Two good paths — pick one today.
You have seen what the tools do. This part gets them running on your own computer or in your browser — and ends with your first line of code and your first error message.
Both paths run every notebook in this course. You can switch anytime — the code is identical.
If setup fights you for more than 30 minutes, switch to Colab, post in the forum, and keep moving. Don't lose the week to an installer.
jupyter lab and press Enterlocalhost:88881 + 1 in the cell, press Shift+Enter2 — you're a Python user nowjupyter lablocalhostAnaconda needs ~5 GB free. Tight on space? Use Colab this week and install later. Snags (antivirus prompts, “command not found”) are in the setup guide on the course page.
1 + 1, press Shift+Enter — doneColab is a genuinely professional tool, not a training-wheels compromise. A session is the Python running behind your notebook right now: when it disconnects, the variables it remembered are gone, but your saved notebook and files stay.
Every course notebook has an "Open in Colab" badge — one click from the course page to a running copy.
A notebook is a sequence of cells (boxes of text or code) run by a live kernel — the Python program behind the page, which remembers everything run so far. To execute or run a cell means to make Python carry it out.
Running cells out of order. The kernel's memory follows execution order, not page order — when confused, "Restart & Run All."
Formatted text, headings, math — the narrative explaining why. Markdown is a
simple way to format text: # Title becomes a heading.
Python that runs on Shift+Enter; the result prints just below.
The Python process holding your variables. Restarting it wipes the slate — which is sometimes exactly what you need.
A variable is a name that holds a value (a piece of data, such as the number 100). Run these three cells in the order 1 → 3 → 2 → 3 and the "same notebook" gives two different answers. The kernel only knows what ran, and when.
Before trusting (or submitting!) a notebook: Kernel → Restart & Run All. If it errors or changes, the notebook was lying to you.
price = 100assignment: put the value 100 in a box labelled price. = means “store”, not “equals”price = price * 2take what is in price now, double it, store it back (* means multiply)print(price)call the print function: show the value on screen — 200 after one run of cell 2, 400 after two| Mode | What it is | Best for |
|---|---|---|
| Notebook (.ipynb) | Cells + narrative + inline output | Exploration, teaching, reports — our default |
| Script (.py) | Plain file run top to bottom | Automation, reusable pipelines (Week 11) |
| REPL | Interactive prompt, one line at a time | Quick checks: "what does this function return?" |
Professionals use all three daily. We start in notebooks because seeing each
step's output accelerates learning. REPL stands for
Read–Evaluate–Print Loop: type one line, see its answer at once. The ending of a
file name (.ipynb, .py) is its extension: it says what kind
of file it is.
Explore in a notebook; when it works and must run repeatedly, promote it to a script.
Don't worry about memorizing syntax — today is about the shape of the work: load, look, count.
"Import pandas. Read the CSV into a table called df. Show the first rows. Count
by region." Python rewards reading code as sentences.
The text in quotes, "psa_population.csv",
is a file path: the address of a file on your computer. A bare file
name means “in the same folder as this notebook”; a path like
data/psa_population.csv means “inside the data folder”.
df = pd.read_csv("psa_population.csv")call pandas’ read_csv function; the table it gives back is stored in dfdf.head()the dot means “do this to df”: show its first 5 rows, so you can see what you loadeddf.shapethe table’s size as a pair: (number of rows, number of columns)df["region"].value_counts()take the region column and count how many rows each region hasAn error is Python stopping because it cannot do what a line asks. The traceback is the report it prints — alarming to look at, but actually a structured report. Read the last line first: error type, then message. The lines above just show where it happened.
NameError — you used a name before defining it (or typo'd it)FileNotFoundError — wrong filename or wrong folderSyntaxError — Python couldn't even parse the lineNameError: …the error type (NameError = a name Python does not know) and the message: pirce was never definedFile "cell 4", line 1where it happened: cell 4, line 1 — go and look at that line^^^^^points at the exact part of the line Python could not handle>>> print(pirce)the line that was run (>>> is the prompt where you type). The fix: spell it priceThe last line of the traceback plus the function's official documentation solves most beginner problems in under five minutes.
Someone has hit your exact error. If 15 minutes of searching fails, post to the course forum: the error, your code, what you tried.
Allowed for explaining errors and concepts. For graded exercises, you must be able to explain every line you submit — and say so if AI helped. Details in the syllabus.
The ladder isn't bureaucracy — each rung builds a skill (error literacy, search craft, question-asking) that employers explicitly test for.
A forum post with a minimal reproducible example — the shortest code that still shows the problem — often answers itself as you write it. (Docs, short for documentation, is the official manual for a language or library.)
Twelve weeks from zero to a working pipeline.
You now have a working setup and have read your first error. This part explains how the rest of the course runs, and what to do before next week.
| Weeks | Theme | Topics |
|---|---|---|
| 1–4 | Python Foundations | This intro · types & functions · data structures & files · OOP and errors |
| 5–7 | NumPy & pandas | Arrays · loading & cleaning tables · grouping, merging, reshaping |
| 8–10 | Visualization & Sources | matplotlib & seaborn · files, SQL & APIs · text & regex |
| 11–12 | Workflows & Project | Git, environments & debugging · integrative project presentations |
Each week: a short lecture deck like this one, a code-along notebook, and one graded exercise. Unfamiliar names in the table are explained in the week they arrive.
A small end-to-end pipeline (a script that takes raw data all the way to a result) on a dataset from your own field — scoped in Week 8, presented in Week 12.
Materials drop Monday; exercises due Sunday night. Within the week, your hours are yours.
Questions go to the course forum so everyone benefits; expect instructor replies within one working day.
Stuck for 30+ minutes? Post what you tried. Debugging in public is a professional skill, not a confession.
Course details: 3 units, graded (GRD), no prerequisites. The course is repeatable for credit per the catalog.
Plan for ~3 hours of contact material plus 3–4 hours of practice weekly. Programming is learned in the fingers.
Every new word from this week. Come back to this slide whenever a word in a later week feels slippery.
clean.py# that Python ignores; it is for human readersprice = 100print("hi")data/grades.csvimport pandas as pdThe Week 1 lab runs in your browser from the course page. Replace each
____, press ▶ Run to see the output, then
Check. It borrows a few ideas from the next two weeks — the lab explains
each one as it goes, and you will meet them properly soon.
You are not expected to have mastered these words yet. The goal this week is to press Run, read what comes back, and fix one small thing at a time.
[28, 31, 29] (Week 3)"Ana" (Week 2)f that fills in values from { }One goal: a working environment with proof.
Install Anaconda (Path A) or sign in to Google Colab (Path B) — links on the course page.
Download week01_hello_data.ipynb from the course page and run every cell
top-to-bottom.
Upload the executed notebook (every cell run, so each result shows underneath it) to the submission portal, and introduce yourself in the forum: field, data you care about, Python experience (zero is fine).
A script is a record that re-runs; clicks are a memory that fades.
Python underneath, NumPy/pandas as workhorses, Jupyter on top as your workbench.
Local or cloud, get one environment running and prove it with an executed notebook.
Everything in Weeks 2–12 assumes the Week 1 checkpoint is done — invest the hour now.
Forum first, 30-minute rule, and remember Colab is always a browser tab away.
McKinney, Python for Data Analysis (wesmckinney.com/book) — by pandas' creator. · VanderPlas, Python Data Science Handbook (jakevdp.github.io/PythonDataScienceHandbook).
docs.python.org (the official tutorial) · pandas.pydata.org/docs · Colab's own intro notebooks at colab.research.google.com.
All links are collected on the course page — no need to copy from slides.
When curious about a function, check its official docs before a random blog post — it's faster than it looks.
Types, control flow, and functions — the grammar of every script you'll ever write. Bring a working environment.
DS 208 · Programming for Data Science