DS 208 · Graduate · Week 1

Programming for Data Science

Why code beats clicks, what the Python data stack looks like, and how to get your machine — or the cloud — ready to compute.

3 units · Distance education (asynchronous online)

Session Map

Three parts, ~one hour each — pause anywhere

Today

Where we're going in Week 1

Words for Today

Ten words you will hear this week

You do not need to memorise these now. Each one is explained again on the slide where it first matters, and they are all collected in a glossary near the end.

script

A file of code — written instructions — that Python runs from the top line to the bottom.

e.g. a file called clean_survey.py

library

A bundle of ready-made code someone else wrote, which you load into your own work.

e.g. pandas (tables), NumPy (numbers), matplotlib (charts)

import

The line that loads a library so your code can use it.

e.g. import pandas as pd — load pandas, call it pd for short

notebook

A document of boxes that mix text and runnable code. Jupyter and Colab show notebooks.

e.g. week01_hello_data.ipynb

cell

One box in a notebook. Shift+Enter runs it; the result appears underneath.

e.g. a cell containing 1 + 1 shows 2

kernel

The Python program running behind a notebook. It remembers everything you have run so far.

e.g. “Restart kernel” wipes that memory clean

variable

A name that holds a value (a piece of data), like a labelled box.

e.g. price = 100 puts 100 in a box labelled price

function

A named piece of work. You call it by writing its name and brackets; inputs go inside the brackets.

e.g. print("hi") calls the function print

error

Python stopping because it cannot do what a line asks. Normal, expected, and fixable.

e.g. a misspelt name gives a NameError

traceback

The report printed with an error. Read the last line first: it names the problem.

e.g. NameError: name 'pirce' is not defined

The Motivation

The spreadsheet ceiling is real

In October 2020, Public Health England lost track of nearly 16,000 positive COVID-19 cases — results vanished because an old Excel format silently truncated rows beyond its limit.

The lesson

Not "spreadsheets are bad" — they're great for small tables. The lesson is that invisible, unrepeatable data handling fails silently at scale.

Clicks leave no trace

A month later, nobody — including you — can say exactly which cells were edited, filtered, or pasted.

Manual work doesn't repeat

New month, new file, same 40 clicks — every repetition is a fresh chance for a new mistake.

Size and complexity ceilings

Millions of rows, nested JSON, joins across ten files, statistical models — the grid runs out of headroom.

A Second Cautionary Tale

The spreadsheet error that shaped austerity policy

Reinhart & Rogoff's influential 2010 paper "Growth in a Time of Debt" argued growth collapses when public debt passes 90% of GDP — and was cited to justify austerity worldwide. In 2013, graduate student Thomas Herndon tried to replicate it and found, among other issues, an Excel formula that silently omitted five countries.

Why replication caught it

The error surfaced only because Herndon requested the actual spreadsheet and re-ran the analysis. No script existed to audit — the analysis lived in cell selections.

Two lessons for us

  • Code makes your analysis auditable — others can find your mistakes, which is a feature, not a threat
  • A graduate student with programming skills checked work that moved economies — that skill set is what this course builds
Part 1

Why Programming for Data Science?

Three superpowers: reproducibility, scale, automation.

You already know how to analyse data by clicking through a spreadsheet. This part shows what you gain when those clicks are written down as code instead.

Superpower 1

A script is an analysis that explains itself

Code is written instructions for the computer. A script is a file of code that runs from top to bottom: a complete, ordered, re-runnable record of every step from raw file to final figure. Hand it to a colleague — or your future self — and the result reappears exactly.

Science angle

Reproducibility is the currency of graduate research. Reviewers and advisers increasingly expect runnable code, not screenshots.

You are not expected to write this yet. Read it like a recipe card: each line is one step, in order, and anyone can follow it again.

# The whole analysis, visible at a glance import pandas as pd df = pd.read_csv("enrollment_2026.csv") df = df.dropna(subset=["program"]) summary = df.groupby("program").size() summary.to_csv("per_program_counts.csv")
  • # The whole analysis…a comment: a note for humans. Python ignores everything after #
  • import pandas as pdload the pandas library and call it pd for short
  • df = pd.read_csv(…)read the CSV file into a table and store it under the name df
  • df.dropna(subset=["program"])throw away rows where the program column is blank
  • df.groupby("program").size()count how many rows there are for each program
  • summary.to_csv(…)save those counts to a new file. No output on screen — the result is the file
Superpowers 2 & 3

Scale and automation come free with code

Task By hand With a script
Clean 1 survey file 20 minutes of careful clicking Write once: ~15 lines
Clean 300 survey files Weeks, with drifting rules The same 15 lines in a loop
Redo everything after a data fix Start over, hope for consistency Re-run: one command
Monthly report, every month A recurring calendar dread Scheduled job, done before coffee
Mindset

What programming actually feels like

The Big Picture

The data science workflow you'll program

Import Tidy Transform Visualize Model Communicate the understanding loop — you'll circle it many times

Import = get the data into Python. Tidy = arrange it so each row is one thing and each column one measurement. Transform = filter, compute and summarise. Visualize = draw it. Model = fit a statistical model. Communicate = explain what you found to someone else.

Quick Check · Part A

Which superpower is missing?

A colleague proudly shows you their analysis: a beautifully formatted Excel workbook where they manually filtered, copied cleaned rows to a new sheet, and built a chart. The numbers are correct. From a programming-for-data-science standpoint, what's the core problem?

A · Excel can't make charts properly
B · The result is probably wrong
C · Nobody — including them — can re-run or audit the steps
D · It should have used a fancier statistical model

Commit to an answer, then click and hold (or tap) to reveal

C — it's a reproducibility problem, not a correctness problem.

The work may be flawless today, but the manual steps are invisible: they can't be audited (Reinhart–Rogoff), can't be re-run when data updates, and can't scale to the next 300 files. The result isn't wrong — it's unrepeatable.

Pause Point

  End of Part A

Good place to stop for today, or stretch for five minutes. Part B is a guided tour of the Python libraries you'll live in all semester.

Part 2

The Python Data Science Ecosystem

One language, a division of labor.

You now know why we write code. This part shows what we write it with: one language, Python, plus a handful of libraries that each do one job.

Language Choice

Why this course uses Python

Language Strengths Our take
Python Readable, huge ecosystem, general-purpose (web, automation, ML) Our vehicle — one language covers the whole workflow
R Statistics-first, superb plotting (ggplot2) Excellent; learn it later if your field favors it
SQL The lingua franca of databases Not a rival — you'll meet it in Week 9
Julia Speed for numerical computing Promising, smaller ecosystem
The Stack

Who does what in the Python data stack

Python — the language (syntax, data structures, files) NumPy fast arrays & math pandas tables: load, clean, reshape matplotlib · seaborn visualization scikit-learn · statsmodels modeling & statistics requests · sqlalchemy data sources (Week 9) Jupyter — the workbench where you drive all of it you work top-down; each layer stands on the one below
60-Second Taste · NumPy

Math on a million numbers at once

NumPy's array is a row of numbers, all of the same kind, that NumPy can do maths on in one go. It replaces the loop: you write the operation once and it applies to every element (every number in the array), very fast. This idea — vectorization, one operation applied to a whole array — is the single biggest mental shift from ordinary programming.

It is like typing a formula once in Excel and having it fill the whole column — without even dragging it down.

Don't memorize — recognize

Today's goal is only to recognize the shape of the code. Week 5 makes you fluent.

import numpy as np # monthly rainfall, mm — one array, no loops rain = np.array([12.4, 8.1, 22.7, 156.3, 301.9, 288.0]) rain * 0.1 # convert all values to cm rain.mean() # 131.6 rain > 100 # [False False False True True True] rain[rain > 100] # just the wet months
  • rain = np.array([…])make an array of six rainfall numbers and name it rain
  • rain * 0.1multiply every number by 0.1 at once — no loop needed
  • rain.mean()the average of all six: about 131.6 mm
  • rain > 100ask “over 100?” of each number; each answer is True (yes) or False (no)
  • rain[rain > 100]keep only the numbers where the answer was True
60-Second Taste · pandas

A spreadsheet you command in sentences

pandas puts labels on NumPy's speed. Its table is called a DataFrame: rows and named columns, like a spreadsheet you control with code. The questions you'd click through in Excel become one readable line each.

A CSV file (“comma-separated values”) is a plain-text table: one row per line, commas between the columns. Excel can save one; pandas can read one.

The 80% library

Weeks 6–7 are pure pandas. Most working data scientists spend most hours exactly here.

import pandas as pd df = pd.read_csv("barangay_health.csv") # the Excel pivot table, as one sentence: df.groupby("barangay")["no_show"].mean() # filter + sort, still readable aloud: df[df["distance_km"] > 4].sort_values("age")
  • df = pd.read_csv("…csv")read the CSV file into a DataFrame called df
  • df.groupby("barangay")["no_show"].mean()for each barangay, the average of the no_show column — a pivot table in one line
  • df[df["distance_km"] > 4]keep only the rows where distance is over 4 km
  • .sort_values("age")then put those rows in order of age
60-Second Taste · matplotlib

From table to picture in three lines

matplotlib draws; seaborn (built on top) makes the common statistical plots pretty by default. Every chart is code — which means every chart is reproducible and revisable.

Charts as arguments

By Week 8 you'll treat a figure the way you treat a paragraph: drafted, revised, and defended.

import matplotlib.pyplot as plt monthly = df.groupby("month")["visits"].sum() monthly.plot(kind="bar") plt.title("Clinic visits by month") plt.show() # the figure appears right in your notebook
  • import matplotlib.pyplot as pltload matplotlib’s drawing tools, nicknamed plt
  • monthly = df.groupby("month")…sum()total visits per month, from a table df loaded earlier
  • monthly.plot(kind="bar")draw those totals as a bar chart
  • plt.title(…) / plt.show()add a title, then display the finished chart
Quick Check · Part B

Right tool, right layer

You have a CSV of 200,000 e-wallet transactions and want the average amount per province. Which layer of the stack does the heavy lifting?

A · Raw Python — write a for-loop over every row
B · NumPy — build a numeric array per province by hand
C · pandas — read_csv, then groupby("province").mean()
D · matplotlib — it computes averages while plotting

Commit to an answer, then click and hold (or tap) to reveal

C — this is the pandas shape: labeled table, grouped summary.

A works but is slow to write and easy to get wrong; B rebuilds what pandas already gives you; D confuses drawing with computing. Matching the task to the layer is half of fluency — and it's learnable.

Pause Point

  End of Part B

Next part is hands-on: have your laptop ready and ~1 hour of focus. You'll leave it with a working environment — this week's actual deliverable.

Part 3

Your Toolkit Setup

Two good paths — pick one today.

You have seen what the tools do. This part gets them running on your own computer or in your browser — and ends with your first line of code and your first error message.

Decision Point

Local install or cloud notebook?

Path A · Step by Step

Installing Anaconda (10–20 minutes)

Path B · Step by Step

Google Colab (3 minutes, honestly)

The Workbench

Anatomy of a Jupyter notebook

A notebook is a sequence of cells (boxes of text or code) run by a live kernel — the Python program behind the page, which remembers everything run so far. To execute or run a cell means to make Python carry it out.

The #1 beginner trap

Running cells out of order. The kernel's memory follows execution order, not page order — when confused, "Restart & Run All."

Markdown cells

Formatted text, headings, math — the narrative explaining why. Markdown is a simple way to format text: # Title becomes a heading.

Code cells

Python that runs on Shift+Enter; the result prints just below.

The kernel

The Python process holding your variables. Restarting it wipes the slate — which is sometimes exactly what you need.

See the Trap Once

Execution order ≠ page order

A variable is a name that holds a value (a piece of data, such as the number 100). Run these three cells in the order 1 → 3 → 2 → 3 and the "same notebook" gives two different answers. The kernel only knows what ran, and when.

The professional reflex

Before trusting (or submitting!) a notebook: Kernel → Restart & Run All. If it errors or changes, the notebook was lying to you.

# cell 1 price = 100 # cell 2 price = price * 2 # run this twice by accident... # cell 3 print(price) # 200? 400? depends on history, # not on what the page shows
  • price = 100assignment: put the value 100 in a box labelled price. = means “store”, not “equals”
  • price = price * 2take what is in price now, double it, store it back (* means multiply)
  • print(price)call the print function: show the value on screen — 200 after one run of cell 2, 400 after two
Three Ways to Run Python

Notebook, script, or REPL — when to use which

Mode What it is Best for
Notebook (.ipynb) Cells + narrative + inline output Exploration, teaching, reports — our default
Script (.py) Plain file run top to bottom Automation, reusable pipelines (Week 11)
REPL Interactive prompt, one line at a time Quick checks: "what does this function return?"
First Contact

Your first cell of real data work

Don't worry about memorizing syntax — today is about the shape of the work: load, look, count.

Read it aloud

"Import pandas. Read the CSV into a table called df. Show the first rows. Count by region." Python rewards reading code as sentences.

The text in quotes, "psa_population.csv", is a file path: the address of a file on your computer. A bare file name means “in the same folder as this notebook”; a path like data/psa_population.csv means “inside the data folder”.

import pandas as pd # population estimates, one row per city/municipality df = pd.read_csv("psa_population.csv") df.head() # peek at the first 5 rows df.shape # (rows, columns) df["region"].value_counts()
  • df = pd.read_csv("psa_population.csv")call pandas’ read_csv function; the table it gives back is stored in df
  • df.head()the dot means “do this to df”: show its first 5 rows, so you can see what you loaded
  • df.shapethe table’s size as a pair: (number of rows, number of columns)
  • df["region"].value_counts()take the region column and count how many rows each region has
Your First Error, Defused

How to read a traceback (bottom-up)

An error is Python stopping because it cannot do what a line asks. The traceback is the report it prints — alarming to look at, but actually a structured report. Read the last line first: error type, then message. The lines above just show where it happened.

This week's whole error vocabulary

  • NameError — you used a name before defining it (or typo'd it)
  • FileNotFoundError — wrong filename or wrong folder
  • SyntaxError — Python couldn't even parse the line
>>> print(pirce) Traceback (most recent call last): File "cell 4", line 1, in <module> print(pirce) ^^^^^ NameError: name 'pirce' is not defined. Did you mean: 'price'? ← it even guesses
  • 1st · last line: NameError: …the error type (NameError = a name Python does not know) and the message: pirce was never defined
  • 2nd · File "cell 4", line 1where it happened: cell 4, line 1 — go and look at that line
  • 3rd · ^^^^^points at the exact part of the line Python could not handle
  • >>> print(pirce)the line that was run (>>> is the prompt where you type). The fix: spell it price
When You're Stuck

The help ladder — climb it in order

Part 4

This Course

Twelve weeks from zero to a working pipeline.

You now have a working setup and have read your first error. This part explains how the rest of the course runs, and what to do before next week.

Roadmap

The 12 weeks at a glance

Weeks Theme Topics
1–4 Python Foundations This intro · types & functions · data structures & files · OOP and errors
5–7 NumPy & pandas Arrays · loading & cleaning tables · grouping, merging, reshaping
8–10 Visualization & Sources matplotlib & seaborn · files, SQL & APIs · text & regex
11–12 Workflows & Project Git, environments & debugging · integrative project presentations
How Async Works Here

Freedom of schedule, not of deadline

Glossary

Week 1 words, one line each

Every new word from this week. Come back to this slide whenever a word in a later week feels slippery.

This Week's Lab

Your first lab: a gentle preview of Weeks 2–3

The Week 1 lab runs in your browser from the course page. Replace each ____, press ▶ Run to see the output, then Check. It borrows a few ideas from the next two weeks — the lab explains each one as it goes, and you will meet them properly soon.

You are not expected to have mastered these words yet. The goal this week is to press Run, read what comes back, and fix one small thing at a time.

Words the lab will introduce

list
an ordered collection of values: [28, 31, 29] (Week 3)
index
an item’s position in a list, counted from 0 (Week 3)
loop
code that repeats a block once per item (Week 2)
string
text inside quotes: "Ana" (Week 2)
f-string
a string starting with f that fills in values from { }
return
how a function hands its answer back (Week 2)
Before Next Week

Your Week 1 checkpoint

One goal: a working environment with proof.

1 · Set up your environment

Install Anaconda (Path A) or sign in to Google Colab (Path B) — links on the course page.

2 · Run the starter notebook

Download week01_hello_data.ipynb from the course page and run every cell top-to-bottom.

3 · Submit your proof

Upload the executed notebook (every cell run, so each result shows underneath it) to the submission portal, and introduce yourself in the forum: field, data you care about, Python experience (zero is fine).

Synthesis

Three things to leave with

Resources

Free references worth bookmarking

Core companions (free online)

McKinney, Python for Data Analysis (wesmckinney.com/book) — by pandas' creator. · VanderPlas, Python Data Science Handbook (jakevdp.github.io/PythonDataScienceHandbook).

Documentation you'll live in

docs.python.org (the official tutorial) · pandas.pydata.org/docs · Colab's own intro notebooks at colab.research.google.com.

Next Week

Python Fundamentals I

Types, control flow, and functions — the grammar of every script you'll ever write. Bring a working environment.

DS 208 · Programming for Data Science