Keeping a program readable once it outgrows a single notebook cell — with classes, error handling, and files that import each other.
Programming for Data Science · University of the Philippines Cebu
Bundle data and the behaviour that belongs with it into one object.
Expect bad input. Catch what you can handle; let real bugs surface.
Split code across files and import it, so projects stay navigable.
So far your code lived in one place. Real projects are many files that lean on each other — today is how.
Write a class, handle an error on purpose, and import your own module.
Last week: lists, dictionaries, and reading files. That fills a notebook fast. Structure is what stops it turning into a wall of tangled cells.
Copy-pasted logic and mile-long cells become impossible to fix or trust.
Group related data and behaviour, guard against bad input, and split into files.
A blueprint (think: cookie cutter) that describes what one kind of thing holds and what it can do.
e.g. class Student: starts the blueprint for a student
One actual thing built from a class — one cookie from the cutter. “Instance” means the same thing.
e.g. ana = Student("2024-001", "Ana") makes one student object
A value stored on an object. You read it with a dot and no brackets.
e.g. ana.name gives 'Ana'
A function that lives inside a class and works on that object’s own data. Called with a dot and ( ).
e.g. ana.average() works out Ana’s mean grade
Inside a class, the name for “this particular object”. Python fills it in for you.
e.g. self.name = name stores the name on this object
The set-up method Python runs by itself every time you create a new object. Said “dunder init” (double underscore).
e.g. Student("2024-001", "Ana") quietly runs __init__
Making a new class from an existing one: it gets everything the parent class has, then adds or changes a little.
e.g. class DiscountedSale(Sale): is a Sale, with a discount
Python’s report that something went wrong while running. It stops the program unless your code handles it.
e.g. ValueError, KeyError, IndexError
The error printout. It shows the route Python took; the last line names the error type and message.
e.g. last line ValueError: invalid literal for int()
Try a risky line; if a named exception happens, run the except block instead of crashing.
e.g. try: int(raw) … except ValueError: score = 0
Stop on purpose and report an exception yourself, with a message you write.
e.g. raise ValueError("out of range: 130")
A .py file of code that other files can import (bring in) and reuse.
e.g. grades.py, then from grades import average
A folder of modules that Python can import from, marked by a file called __init__.py.
e.g. a src/ folder holding load.py and clean.py
Data and behaviour, bundled into one thing.
You already know that a function groups steps (week 2) and a dictionary groups values (week 3). A class does both at once: it keeps the values and the functions that use them together.
A student has a number, a name, and grades. Passing three separate variables (named boxes) everywhere is fragile. A class is a blueprint for one kind of thing. Each thing built from it is an object (also called an instance) that keeps those values together — with its own methods (functions that belong to it).
A class is a cookie cutter; each object is one cookie. You make the cutter once, and every cookie has the same shape but its own toppings — Ana’s grades are not Ben’s grades.
Groups behaviour: give it inputs, get an output.
Groups data + behaviour: the values and the actions that belong with them.
__init__ sets up each new object__init__ (“dunder init”) is the set-up method:
it runs by itself when you create an object. self is
this particular object, and an attribute is a value stored on it
with a dot (self.name). The attributes you set are that object’s own.
class Student:start a blueprint called Student (class names start with a capital letter by habit)def __init__(self, sid, name):the set-up method; sid and name are the two inputs you will giveself.sid = sidstore the input on this object as an attribute called sid (same for name)self.grades = []every new student starts with their own empty list of gradess = Student("2024-001", "Ana")make one object; Python runs __init__ with self = the new students.nameread an attribute with a dot: "Ana" is the name stored on sselfA method is a function written inside a class. Its first
parameter is self, so it can reach the object’s own attributes.
Because the grades live on the object, the object can compute its own average.
No need to pass the data back in — it is already there.
# ...__init__ as before...the set-up code from the previous slide is still there; it is just not repeateddef average(self):a method: no extra inputs, only self — the student it is asked aboutsum(self.grades) / len(self.grades)add this student’s grades and divide by how many there ares.average()call the method with a dot and ( ); 86.33 is Ana’s mean, 259 ÷ 3 = 86.333… shown roundedThe class is written once. Every Student(...) call makes a fresh
object with independent attributes.
One cookie cutter, many cookies. Student is written once; Ana, Ben and Cy are three separate objects, so changing Ben’s grades never touches Ana’s.
A pandas DataFrame (the table you will control with code in week 6) is an object too — data plus methods like
.mean().
__init__ runs the moment you make an objectIt is not magic and it is not a constructor in the C++ sense (if you have never met C++, ignore that: it simply means __init__ is ordinary Python). It is a plain method Python calls for you, with the new object as self.
Making an object is like filling in a new form. Sale("kape", 15, 3) hands Python three answers, and __init__ copies each one onto the new form, which it calls self.
self.item = itemcopy the input item onto this new object (“attach to THIS object”)s = Sale("kape", 15, 3)create one Sale: kape (coffee), price 15, quantity 3print(s.item, s.price)print two attributes; kape 15 is the two values with a space between themself is just the first argumentYou never pass it. Sale("kape", 15, 3) becomes __init__(s, "kape", 15, 3) behind the scenes.
An attribute is data the object carries. A method is a function that can reach that data through self.
An attribute is a noun — what the sale has (a price). A method is a verb — what the sale can do (work out its total).
self.item, self.price, self.qty = item, price, qtyset three attributes in one line: each name on the left gets the value in the same place on the rightdef total(self):a method that works out price × quantity for this sales.pricean attribute is stored data, so no brackets: 15s.total()a method is an action, so it needs ( ) to run: 15 × 3 = 45Print an object you wrote and Python falls back to its type and memory address. Useless in a loop, useless in a debugger.
0x103400e00); tells you nothing about the dataprint() prefers it when a class has both (the lab uses it)f before the quote; each {name} inside is replaced by its value: f"{2+3} cups" → '5 cups'<__main__.Sale object at 0x…>the default: a Sale object from your main program, at some memory addressdef __repr__(self):indented, so it goes inside class Sale; Python calls it whenever it has to show the object{self.item!r}fill in the item; !r shows it the way you would type it, in quotes: 'kape'Sale('kape', 15, 3)now printing shows what is inside the objectprint([s1, s2]) now shows both sales instead of two addresses — which is when you actually need it.
@dataclass writes the boilerplateIf a class is mostly "hold these three fields", Python will generate __init__, __repr__ and == for you.
@ just above a class or function; it adds features to what is below__init__, __repr__ and == Python writes for youfrom dataclasses import dataclassbring the dataclass tool in from Python’s built-in dataclasses module@dataclassthe decorator: it rewrites the class below, adding set-up, printing and comparingitem: stra field called item that should hold a string (a hint: Python does not check it)Sale(item='kape', price=15, qty=3)the automatic __repr__: every field with its valueTrue== now compares the fields, so two sales with equal fields count as equalTwo plain objects with identical fields are not equal by default — Python compares identity. A dataclass compares the fields.
In a notebook, build it from scratch. Three fields, one method, one repr.
item, price, qty)__repr__ (two slides back)[x * 2 for x in [1, 2, 3]] gives [2, 4, 6]Sale dataclass with item, price, qty.total() returning price × qty.A dataclass prints its fields, so the list is readable without writing __repr__ yourself.
Inheritance builds a new class on an existing one. A discounted sale is a sale. Subclass it, override only what differs, and call super() for the rest.
Sale) and the new one built on it (DiscountedSale)Inheritance is a recipe variation: “chocolate cake” starts from the plain cake recipe and only changes the flavour. It passes the is-a test: a chocolate cake is a cake.
class DiscountedSale(Sale):a new class built on Sale (the parent goes in the brackets); it gets all of Sale’s methodssuper().__init__(item, price, qty)let the parent’s set-up store the three usual attributes…self.pct = pct…then add the one extra attribute only a discounted sale hasdef total(self):override: this total replaces the parent’s for discounted salessuper().total() * (1 - self.pct/100)the parent’s total (15 × 4 = 60) with 25% off = 45Sale('kape', total=60)printed by a __repr__ on Sale (not shown) that writes the class name and total()Deep hierarchies are hard to follow. If the child does not genuinely pass an is-a test, prefer composition — hold the object instead.
Put a mutable default on the class and all instances point at the same list. One cart fills another.
class: one copy, shared by every objectself in __init__: each object gets its ownitems = []directly under class: made once, shared by every Cartself.items.append(x)the object has no list of its own, so this adds to the shared one['kape'] ['kape']a and b print the same list: b “has” kape though you only added it to adef __init__(self): self.items = []the fix: each new Cart gets its own empty list, so b stays []Assignments directly under class run once, when the class is defined. Assignments in __init__ run per object.
The empty list in the signature is created when the def line runs — not on each call. It then persists between calls.
def line: the function’s name and its parametersbasket=[]is Nonebasket=[]the [] is made once, when def runs, and reused by every calladd_sale("load")the second call gets that same list, still holding kape: ['kape', 'load']if basket is None: basket = []the fix: no basket given → build a fresh list inside, on every callNever use a list, dict or set as a default. Use None and build it inside — that runs on every call.
Next: when things go wrong, and how Python tells you.
Real data breaks assumptions — plan for it.
You met your first traceback in week 1 and wrote if checks in week 2. This part adds how to read an error calmly, catch the ones you expect, and raise your own.
When Python can't continue, it raises an exception — a report of what went wrong — and prints a traceback, the error printout. The last line names the problem; read it bottom-up.
An exception is like a smoke alarm: annoying, but it tells you something real, and the label on it (IndexError, ValueError) says which room to check.
IndexError: list index out of rangeread it as type: message. IndexError = asked for a position that does not exist; nums only has positions 0 and 1ValueError: invalid literal for int()ValueError = the right type (text) but a value int() cannot turn into a whole numberThe last line is what actually broke. The lines above are the route Python took to get there — newest call last.
Here: load_price hit the string "n/a". That is your bug, on line 2.
Traceback (most recent call last):the heading: the calls (functions running because another line asked) are listed oldest first, newest lastFile "sales.py", line 8, in <module>where it started: line 8 of sales.py, at the top level of the file (<module>)File "sales.py", line 2, in load_pricethe deepest call: line 2, inside the function load_price — check this line firstValueError: … 'n/a'the error type and the bad value: the text 'n/a' cannot become a whole number (base 10 = ordinary digits)| Exception | What Python said | Usually means |
|---|---|---|
ValueError | invalid literal for int() with base 10: 'n/a' | Right type, impossible value |
TypeError | can only concatenate str (not "int") to str | Wrong type entirely (joining, or “concatenating”, text with a number) |
KeyError | 'price' | That key is not in the dict |
IndexError | list index out of range | Past the end of the list |
FileNotFoundError | [Errno 2] No such file or directory | Wrong path, or wrong working dir (the folder Python is running in) |
ZeroDivisionError | division by zero | An empty group you averaged |
AttributeError | 'str' object has no attribute 'push' | Wrong method for that type |
The exception name is the diagnosis and the message is the symptom. KeyError: 'price' reads “I looked for the key price and it was not there.”
try the risky line, except the falloutWrap the line that might fail in a try block. If it raises the exception you named, the except block runs instead of crashing.
try/except is a safety net under one trick: if the trick fails in the way you expected, you land softly and the show goes on.
raw = input("Score: ")input() shows the prompt and waits for typing; it always gives back text. Here the user typed "ninety"try:attempt the indented line(s) belowscore = int(raw)the risky line: int("ninety") raises a ValueErrorexcept ValueError:if exactly that error happens, jump here instead of crashingscore = 0the fallback (a sensible replacement): the program carries on with 0else and finally earn their placeMost people learn try/except and stop. The other two say when code runs, which is the part that prevents bugs:
else runs only when nothing went wrong, finally runs every time.
# parse("15")picture these lines inside a function parse(v); v is the text to convertelse:good input: try worked, so else prints parsed 15finally:runs in both cases — the place for clean-up{v!r}bad input: except runs, and !r keeps the quotes so you can see it was textelseCode in try that cannot raise still gets its exceptions caught by accident. else keeps the guarded part honest.
except hides the thing you neededA bare except is except: with no error name. It catches everything — including typos in your own code, and Ctrl-C (the keys that stop a running program). You get "something went wrong" and no way to find out what.
except:nothing after it: catches every error, even a misspelled variable namesomething went wrongthe output says nothing about what or whereexcept ValueError as e:catch only ValueError, and give the caught error the name e{e}printing e shows Python’s own message, so you still see the bad value 'n/a'An exception you did not predict should crash loudly in development. Silencing it just moves the failure somewhere harder to find.
| Exception | Happens when | Example |
|---|---|---|
ValueError | right type, wrong value | int("abc") |
KeyError | dict key missing | pop["Manila"] |
IndexError | list position too big | nums[99] |
FileNotFoundError | file isn't there | open("nope.txt") |
ZeroDivisionError | dividing by zero | total / 0 |
TypeError | wrong type entirely | "a" + 1 |
Name the specific one you expect after except. Catching the exact exception means real bugs
still surface loudly.
Match the exception to the failure — never a blanket catch-all.
Don't let a nonsense value slip downstream. Check it at the door and
raise a ValueError with a message that says what went wrong.
To raise is to stop on purpose and report an exception yourself.
if not 0 <= g <= 100:“if g is not between 0 and 100” — Python allows chained comparisons like in mathsraise ValueError(f"out of range: {g}")stop here and report a ValueError with our own messagereturn gonly reached when the grade is fineValueError: out of range: 130the last line of the traceback (only that line is shown): the type, then our messageA custom exception is your own error type: a class that inherits from a built-in one. Subclass the closest built-in. Callers can then catch your failure specifically, or fall back to the general one.
class PriceError(ValueError):a new kind of error, built on ValueError — Part A’s inheritance, used for errors"""Raised when…"""a docstring: the class’s one-line description; the class needs no other coderaise PriceError(f"…{p}")raise your own type, with a message that includes the bad priceexcept PriceError as e: print(e)catch it by name; printing e gives price cannot be negative: -5Because PriceError subclasses ValueError, code that only knows about ValueError still works.
with closes things even when you crashA context manager is anything you can use with with: it guarantees the cleanup runs — no finally, no forgotten .close().
with is a library book with automatic return: when you leave the indented block, the book goes back — even if you tripped on the way out.
with open("sales.csv") as f:open the file and call it f for the indented block; with promises to close it afterwardsrows = f.readlines()read every line of the file into a list (week 3)print(f.closed)outside the block: True means the file is already closedDatabase connections, locks, timers and matplotlib figures all use the same pattern. Learn it once here.
You have a sensible fallback (a replacement plan) — skip a bad row, use a default, retry.
It signals a bug in your logic. Hiding it just delays the pain.
A bare except: that swallows everything — it buries typos too.
The DS 227 echo again: the default is a choice. Skipping, defaulting, or flagging a bad value is a decision you make on purpose.
Catch narrow, handle honestly, log (keep a written note of) what you skipped.
Errors are easier to learn by causing them. Run each, and name the exception before you look.
int("12.5")[1,2,3][3]{"a":1}["b"]open("nope.csv")ValueError · IndexError · KeyError · FileNotFoundError. The last one names the file it could not find: check the spelling, and the working directory (the folder Python is running in).
Your script dies with KeyError: 'price'. Where do you look first?
B — bottom-up.
The final line names the failure; the file directly above it is where the bad code lives. The top of a traceback is just where your program started.
Reading ages["Rizal"] from a dict
that has no "Rizal" key raises which exception?
ValueErrorIndexErrorKeyErrorTypeErrorC — KeyError.
Missing keys raise
KeyError; missing list positions raise IndexError.
Use .get() when a gap is expected.
Key vs index is the classic mix-up. The container decides the error.
Dict → KeyError. List → IndexError.
Back to give all this code a tidy home across files.
From one long notebook to files that import each other.
You can already write functions and classes in one notebook. This part moves them into files (modules) that any notebook can import — recipes kept in a shared binder instead of on sticky notes.
.pyA module is a .py file of Python code. Put your functions in grades.py. Any notebook or script can then
import them (bring them in by name) — one definition, used everywhere, fixed in one place.
A module is a recipe card in a shared binder: write it once, every cook (notebook) can use it, and fixing a mistake on the card fixes it for everyone.
# grades.pya plain file called grades.py, saved in the same folder as the notebook: this is the module# analysis.ipynbyour Jupyter notebook (.ipynb), the file you work infrom grades import averagefind grades.py and bring the name average into this notebookaverage([88, 92])works as if written here; 90.0 is a float because / always gives a decimal| Form | You then write | Use when |
|---|---|---|
import grades | grades.average(x) | you want the source obvious |
from grades import average | average(x) | you use one thing a lot |
import numpy as np | np.mean(x) | the community's short alias |
An alias is a short nickname you choose with as.
Aliases like np and pd are conventions — everyone
recognises them. Stick to the standard ones.
from grades import * (* = “everything”) — it hides where names came from.
All three work. The middle one is what most real code uses, because the reader can see where a name came from.
It can silently overwrite a name you already had, and a reader cannot tell which module a function belongs to.
if __name__ == "__main__" is forA file has two lives: something you run, and something you import. This line separates them.
"__main__" when you run that file directly, its own name ("sales") when another file imports itA file can be run like a program or borrowed like a toolbox. The guard marks the lines that should happen only when it is run as a program.
def total(rows): ...the ... just means “body left out to save space”if __name__ == "__main__":true only when you type python sales.py to run this file itselffrom sales import totalimporting sets __name__ to "sales", so the indented print is skippedImporting the module would execute your test code, read files and print output — every time, from anywhere.
Separate raw data, code, and notebooks. A README says what the
project is; requirements.txt lists what to install.
data/ ends in / to show it is a folder.md = Markdown, plain text with simple formatting)pandasIf a cleaning step is wrong you rerun it. If you edited the raw file by hand, that is gone forever.
Anything you need twice moves into src/. A notebook
is for looking, not for the logic you depend on.
What this is, how to run it. Write it on day one, while you still remember.
src/ is a package: a folder of modules. The (often empty) __init__.py file tells Python it may import from it__main__ guardCode under if __name__ == "__main__": runs only when the file is
launched directly — not when another file imports it. So importing never triggers side
effects.
def main():a function holding what the script should do when runif __name__ == "__main__":call main() only when grades.py is run directly; importing it runs nothingAfter from stats import mean, how do
you call it?
stats.mean(x)mean(x)import mean(x)stats(mean, x)B — mean(x).
The from ... import name form
pulls the name in directly, so you use it bare. You'd write stats.mean(x)
only after a plain import stats.
How you import decides how you call. Match the two in your head as you type.
from m import f → call f() bare.
You'll write a small Student class, add methods, make it print
readably, build a subclass, catch a bad value with try/except, and raise a clear
error on a bad grade. ~45 minutes, all in the browser.
__init__, self, methods, __str__, inheritance,
try/except/finally, raise, and a custom exception.
Move your Student class into student.py in a real notebook and import it back.
Data + behaviour in one object. __init__ builds it; self is
this one.
Catch the specific exception you expect; raise on bad input; never
swallow all.
Reusable code in modules, imported cleanly, in a layout others can read.
One sentence: group what belongs together, expect what can go wrong, and give every piece a home.
Structure a project instead of piling it into one cell.
Attributes hold, methods do, __init__ sets up. Add
__repr__ or use @dataclass so it prints.
Last line = what broke. The file above it = where. That alone solves most bugs you will hit this term.
Name the exception you expect. Let the ones you did not predict crash loudly while you can still find them.
Into a module, behind if __name__ == "__main__", in
a layout someone else could navigate.
ana.name, no bracketsana.average()f"…" text where each {name} is replaced by its value__init__, __repr__ and == for you[x * 2 for x in nums]else: only when no error; finally: always, for clean-upexcept: with no name — catches everything, hides bugs.py file whose code other files importas__init__.py fileif __name__ == "__main__": — runs only when the file is run directlyFinish the Week 4 lab and submit it. Make sure you can write a class and a
try/except from memory.
Python Tutorial §8 (errors and exceptions) and §9.1–9.4 (classes) —
docs.python.org/3/tutorial.
Everything here is linked on the course page beside this deck.
Next week we leave loops behind for whole-array maths with NumPy.
Arrays that do maths on thousands of numbers at once — faster and shorter than any loop you'd write by hand.
DS 208 · Programming for Data Science