Holding many values at once — lists, tuples, sets and dictionaries — then reading and writing real files, up to and including CSV.
Programming for Data Science · University of the Philippines Cebu
Order, slices, sorting — and the aliasing trap
Fixed records and uniqueness
Lookups, tallies — and nested data
Read, clean, compute, write back
Every dataset you'll ever load — a CSV, a JSON file, a database row — is one of these shapes underneath.
By the end you can read a CSV of grades and write back a per-section summary — a complete, real program.
Last week: a value, a choice, a function. But real data comes in bulk — thousands of rows. You need containers, and a way to load them.
A container is a value that holds other values — one variable, many items inside.
Lists and dictionaries do most of the work; tuples and sets fill two specific gaps.
Getting data in — from a file — so you stop typing it into cells.
Last week each variable held one value: a number, a string, a boolean. Today’s words are about containers (values that hold other values) and about reading data from files.
An ordered collection of values in square brackets. You can change it later.
e.g. [3, 5, 4]
An item’s position number, counted from 0 — like seat numbers that start at zero.
e.g. in ["a", "b"], "b" is at index 1
A run of items taken by position: from the start index up to, not including, the stop.
e.g. [10, 20, 30][0:2] is [10, 20]
A function that belongs to a value, called with a dot after the value.
e.g. names.append("Ana")
An ordered collection in round brackets that cannot be changed once made.
e.g. a coordinate (10.3, 123.9)
An unordered collection in curly brackets that keeps only one of each value.
e.g. set(["yes", "no", "yes"]) is {'yes', 'no'}
A collection of key → value pairs, like a phone book: look up a name, get a number.
e.g. {"Cebu": 964000, "Davao": 1777000}
The label you look a value up by in a dictionary. Each key appears once.
e.g. pop["Cebu"] uses the key "Cebu"
The address of a file on disk. A bare name means “in the current folder”.
e.g. "grades.csv" or "data/grades.csv"
A plain-text table: one row per line, commas between the columns, often a header row first.
e.g. name,score then Ana,88
An ordered shelf you reach by position.
You met lists briefly in Week 2’s loops. This part shows how to reach into one by position, cut pieces out of it, sort it — and the one trap that surprises everybody.
0A list is an ordered collection of values in square
brackets. Each item has an index: its position number. The first item is
[0], not [1]. Negative indexes count from the end, so [-1]
is always the last item.
cities = [...]a list of four strings, stored under the name citiescities[0]square brackets after a list mean “the item at this index”: "Cebu"cities[-1]count from the end: the last item, "Baguio"len(cities)how many items: 4. So the last index is 4 - 1 = 3A slice like cities[1:3] takes from the
start index up to but not including the stop index — here, Davao and Iloilo.
The list is a row of seats
numbered from 0. cities[2] means “whoever sits in seat 2”;
cities[1:3] means “everyone from seat 1 up to, but not including, seat 3”.
[1:3] gives two items, not three. The stop is exclusive — memorise it now.
| Do | Code | Result on [3, 5] |
|---|---|---|
| Add to the end | nums.append(9) | [3, 5, 9] |
| How many | len(nums) | 2 |
| Is it in there? | 5 in nums | True |
| Sum them | sum(nums) | 8 |
| Walk each one | for n in nums: | 3, then 5 |
These five cover the vast majority of list work. Notice for n in
nums gives you the values, not the positions. .append is
a method: a function that belongs to a value, called with a dot after it.
5 in nums asks “is 5 one of the items?” and answers True or
False.
Last week's running-total loop was exactly this: walk a list, accumulate.
b = a does not copy the list — it gives the
same shelf a second name, an alias. Change it through either name and
"both" change, because there is only one.
b = a is like giving a friend a second key to your house,
not building them a new house. If they move the furniture, your house changes too.
.copy() builds the second house.
Want an independent copy? Say so: .copy() (or list(a)).
sorted() returns; .sort() rearrangesTwo tools, one crucial difference: sorted(x) hands back a
new sorted list; x.sort() reorders the original
in place (it changes the list itself rather than making a new one) and
returns None.
x = x.sort() sets x to None. Use one tool or the
other, never both at once.
sorted(temps)a new list, smallest to largest; temps itself is untouchedreverse=Truea keyword argument: sort largest first instead[:2]slice with no start = “from the beginning”: the first two items, i.e. the top twosplit cuts; join gluesA line of text becomes a list of pieces with .split(); a list
becomes one string with .join(). This pair is the bridge between files and
containers — remember it for Part D.
split is called on the string; join on the
separator (the text placed between pieces, here "," or
" | "). Everyone trips on that once. The quotes in '964000' say it is
still a string.
Given nums = [10, 20, 30, 40],
what does nums[1:3] return?
[10, 20, 30][20, 30][20, 30, 40][10, 20]B — [20, 30].
Start at index 1 (20), stop
before index 3 (40). The stop index is never included — the
slice has 3 − 1 = 2 items.
Slice length is stop − start. That arithmetic saves you every
time.
Stop is exclusive. Always.
Two smaller containers that each do one thing perfectly.
You now know lists: ordered and changeable. Each container in this part gives up one list feature on purpose — a tuple cannot change; a set has no order and no duplicates.
A tuple uses round brackets instead of square. Index and slice exactly like a list — but it is immutable (cannot be changed after it is made), so any attempt to modify it fails loudly. That rigidity is the feature.
TypeError: this type (a tuple) does not allow that operationpoint[0] = …Fixed-shape records: a coordinate, an (R, G, B) colour, a (name, score) pair. Position means something.
Unpacking: assign a tuple to comma-separated names and
Python hands out the values in order — the first to lat, the second to
lon. You will use this all through Part C: a dictionary’s
.items() hands you (key, value) tuples. (tally below is
the vote-count dictionary built in Part C.)
Two names need exactly two values — otherwise
ValueError: too many/not enough values to unpack. A ValueError means
the type was right but the value did not fit (here: the wrong number of items).
A set: curly brackets, no duplicates, no positions. Two jobs it does better than anything else: dedup (removing duplicates) and fast membership tests (answering “is this value in there?”).
set(votes)make a set from the list: each answer kept once. No index, so the order may varylen(set(votes))how many different answers: 2"maybe" in set(votes)is "maybe" one of them? False"How many distinct cities are in this column?" —
len(set(column)). One line.
| Question | Operator | {1,2,3} vs {3,4} |
|---|---|---|
| In either group (union) | A | B | {1, 2, 3, 4} |
| In both groups (intersection) | A & B | {3} |
| In A but not B (difference) | A - B | {1, 2} |
"Which students submitted the lab but not the quiz?" is
submitted_lab - submitted_quiz. Data questions are set questions,
constantly.
Membership (in) on a set is near-instant regardless of size; on a long
list it's a full scan. Picture two overlapping circles: union is everything
in either circle, intersection the overlap, difference the
part of A outside B.
| Container | Looks like | Ordered? | Changeable? | Reach for it when… |
|---|---|---|---|---|
| list | [1, 2, 2] | Yes | Yes | the items, in order |
| tuple | (10.3, 123.9) | Yes | No | a fixed-shape record |
| set | {1, 2} | No | Yes | uniqueness & membership |
| dict | {"a": 1} | Insertion | Yes | look things up by name |
Ninety percent of programs use lists and dicts; tuples and sets are the right tool often enough that not knowing them costs you.
A DataFrame (pandas’ table, Week 6) column behaves like a list; a row like a dict; its row labels like a set. These four are the atoms. “Changeable” is also called mutable; “Insertion” means items stay in the order you added them.
A column of 500 barangay names has duplicates. Which expression counts the distinct barangays?
len(names)len(set(names))set(len(names))sorted(names)[0]B — len(set(names)).
set() collapses duplicates;
len() counts what's left. A gives 500 (rows, not barangays); C is
backwards and errors; D is just the alphabetically first name.
Read compound expressions inside-out: set(names) first, then
len(...) of that.
Distinct count = len(set(x)). You'll type it weekly.
Back for dictionaries, nested data, and real files.
A labelled drawer you reach by key, not by position.
Lists and tuples find things by position. This part adds the dictionary, which finds things by name — and then puts containers inside containers, which is how real datasets look.
When "the third item" is meaningless but "Davao's population" is exactly what you want, use a dictionary. Each entry is a key : value pair: the key is the label you look up, the value is what you get back.
A dictionary is a phone book: you look up a name (the key) and get a number (the value). Nobody asks for “the 3rd entry in the phone book”.
pop = { … }curly brackets with key: value pairs make a dictionary964_000the underscores are only for your eyes: Python reads 964000pop["Davao"]look up the value stored under the key "Davao": 1777000pop["Baguio"] = 366_000a new key: adds the pair. (An existing key would be overwritten.)| List | Dictionary | |
|---|---|---|
| Reach an item by | position: x[0] | key: x["Cebu"] |
| Order matters? | Yes — it's a sequence | Insertion order kept, rarely used |
| Best when | "the items, in order" | "look this up by name" |
| Missing item | IndexError | KeyError (or .get) |
Reach for a list when order is the point; a dictionary when you'll look things up by a meaningful name.
A spreadsheet row is often a dict; a whole column is often a list.
[] crashes; .get() copesAsking for a key that isn't there raises KeyError. When a gap
is expected — and in real data it always is — .get() hands back a default
instead.
KeyError: 'Manila' — the dictionary has no key 'Manila'; the message is the missing key"manila" and "Manila" are different keys.get(key, default)None) instead of stoppingA dictionary maps each distinct item to a running count. This one pattern underlies word counts, vote counts, category frequencies — everything.
tally = {}start with an empty dictionaryfor v in votes:one pass per vote; v is "yes", then "no"…tally.get(v, 0) + 1the count so far for this vote (0 if it is new), plus onetally[v] = …store the new count under that key{'yes': 3, 'no': 1}the result: three yes votes, one no| Loop | Gives you | Use when |
|---|---|---|
for k in d: | each key | you only need the labels |
for v in d.values(): | each value | you only need the numbers |
for k, v in d.items(): | both at once | you need the pair — the common case |
.items() is the one you'll reach for most: it hands you the key
and its value together — as a tuple, unpacked into k, v.
Plain for k in d gives keys, not values — a classic surprise.
You loop for x in scores: over a
dictionary scores. What does x hold each time?
(key, value) pairsB — the keys.
Looping a dict directly gives its
keys. For the values use .values(); for both use
.items(). Assuming you get values is a very common bug.
When you need the value, say so: d[k] inside the loop, or
.items().
Bare loop over a dict → keys.
Nested data is a container inside a container. List of dicts = one dict per row. Dict of lists = one list per group or column. Every CSV, every JSON response (a common text format for web data, Week 9), every DataFrame is one of these two shapes.
DS 227's scraper builds a list of dicts, then pd.DataFrame(rows). Same
shape, same reason.
Each bracket peels one layer. students[0] is a dict;
["name"] reaches into it. Read chains left to right, one hop at a time.
Like a postal address: first the building (students[0]),
then the room inside it (["name"]).
Lost in a nest? Print one layer at a time: print(x), then
print(x[0]), then print(x[0]["name"]).
Loop over the rows; each pass holds one dict. Every filter, tally, and "who's top?" question is this loop wearing different clothes. (The grey lines show what the variables hold afterwards.)
for s in students:s is one student’s dict each time rounds["score"] >= 80that student’s score — at least 80?top = students[0]start by assuming the first student is the besttop = sfound a higher score: remember this student insteadThe tally pattern's big sibling: instead of counting per key, you collect per key. First time you see a key, start its list; then append.
When pandas gives you df.groupby("section"), this loop is what it's doing
underneath.
if sec not in by_section:first time we meet this section? (in on a dict checks its keys)by_section[sec] = []then give it an empty list to fillby_section[sec].append(…)add this student’s score to their section’s listReading data from disk instead of typing it into a cell.
So far every value was typed into the code. You now have containers to hold data; this part fills them from files on disk, and writes results back out.
A file is named data saved on disk; its
file path is its address. "hours.txt" alone means
“in the current folder”; "data/hours.txt" means “inside the
data folder”. Loop over an open file and you get one line per pass. The
with block closes the file for you when it ends — even if
something errors.
open("hours.txt")open the file for reading; a wrong path gives FileNotFoundErrorwith … as f:call the open file f inside the indented block; close it at the endfor line in f:one line of text per pass — always a stringrepr(line)show the string with its hidden characters: \n is the end-of-line mark\nEach line carries an invisible newline character,
written \n: the mark that says “the line ends here”. repr()
reveals it. Leave it in and comparisons and conversions misbehave in ways you won't see.
.strip() removes whitespace (spaces, tabs, newlines) from both ends.
.strip() every line as you read it. Remove whitespace before you trust the
text.
Strip each line, convert it to a number, accumulate. This tiny loop is the skeleton of every data-loading script you'll ever write.
total = count = 0start both at 0 in one lineif not line: continuean empty string is falsy; continue = skip the rest of this pass, go to the next linetotal += int(line)turn the text into a number and add it (+= means total = total + …)print(total / count)the average of the non-blank linesif not line: continue — an empty string is "falsy," so this skips them.
.strip() removes leading/trailing spaces, tabs and the newline.
Is a 0 a real reading or missing data? The file can't tell you — you
decide.
Loading a file is never just reading it — it's a stream of small decisions, exactly like DS 227's Data Preparation phase.
"The default is a choice" applies here too: skip, keep, or flag each oddity on purpose.
"w" — and its dangerOpen with a mode string, which says what you plan to
do: "r" reads (the default), "w" erases and rewrites
the file; "a" appends to the end. Newlines are yours to add — write()
doesn't.
open("data.txt", "w") empties the file the instant it runs —
before any write. Never open your raw data in "w".
split returnsReal data files have several values per line, separated by commas: a
CSV file (comma-separated values). The first line is usually the
header row — the column names. Part A's .split(",")
turns each line into a list — a row.
split(",") breaks on real CSVsThe moment a value itself contains a comma, naive splitting shreds the
row. CSV’s quoting rule: a value containing a comma is wrapped in double quotes. The
csv module — a file of ready-made code that comes
with Python, loaded with import csv — knows all the rules.
next(...) takes the first row the reader produces.
split is fine for files you wrote. Anyone else's CSV: use the
csv module.
csv.reader: each row arrives splitWrap the open file; loop rows as lists. next(reader) pulls
the header row off the top before the loop.
reader = csv.reader(f)wrap the open file: each row will come out as a list of stringsheader = next(reader)take the first row (the column names) off the topfor row in reader:the remaining rows, one list per passDictReader: rows you read by nameIt uses the header row as keys, so each row is a dict —
row["score"] instead of row[2]. This is the list-of-dicts shape
from Part C, straight off the disk.
'88', not 88. CSV knows nothing about types —
int(row["score"]) is on you.
list(csv.DictReader(f))read every row as a dict (header names as keys) and collect them in a listlen(rows) → 5five data rows; the header is not countedrows[0]the first row: a dict, so rows[0]["name"] is 'Ana'csv.writer handles the commas and quoting for you — one
writerow per line, header first. w.writerow([...]) takes a list and
writes it as one comma-separated line. The grey block shows what ends up in the file.
newline=""Required quirk when writing CSV — without it Windows gets blank lines between rows. Just always include it.
After rows =
list(csv.DictReader(f)), what is rows[0]["score"] + 2?
90'882'TypeError'88 + 2'C — TypeError: can't add
str and int.
CSV values are always strings —
rows[0]["score"] is '88'. You must convert first:
int(rows[0]["score"]) + 2 → 90. (B would be the answer
for + "2" — string concatenation.)
Files hand you text. Numbers exist only after you convert.
Convert at the moment of reading — int(...)/float(...) right
in the loop.
Read a grades CSV → group scores by section → write a summary CSV. Containers, nesting, files — in twelve lines.
DictReader loads rows; the grouping move collects scores per
section; csv.writer saves the answer. Spot all three patterns.
rows = list(csv.DictReader(f))1 · read: every row as a dictby_sec[sec].append(int(…))2 · group: each score, converted to a number, into its section’s listfor sec in sorted(by_sec):3 · write: the sections in order A, B (sorting a dict sorts its keys)round(… , 1)the mean, rounded to 1 decimal placeThis ran for real while building the deck. Input: 5 students, 2 sections. Output: one clean summary another program (or Excel) can open.
Section A: (88 + 91 + 67) / 3 = 82.0. Never trust a pipeline you haven't spot-checked once by hand.
You'll slice and sort a list, unpack a tuple, count distinct values with a set, tally votes and group scores in a dictionary, write a small file and read it back, then read a CSV and average it per section — the full load-and-summarise loop. ~45 minutes.
How it works: replace each ____, press
▶ Run, then Check. Each part starts with a short
reminder of the words it uses.
Indexing and slicing, .get() lookups, the tally pattern, and
with open(...).
Re-run tonight's capstone on your own CSV — change a score, watch the summary move.
Stretch: add a count column per section.
Slices stop early; b = a aliases; tuples freeze a record's shape.
.get() for safety; tally and grouping patterns;
len(set(x)) for distinct counts.
List of dicts = rows. Dict of lists = groups. Chain brackets outside → in.
with open, .strip(), csv.DictReader in,
csv.writer out.
One sentence: load rows into containers, clean as you go, group with a dict, and write the answer back out.
Build a complete file-in, file-out data program — no pandas required (yet).
Every new word from today. pandas (Weeks 6–7) is built on all of them.
[3, 5, 4]-1 is the last itemx[start:stop]: from start up to, not including, stopx.append(9)b = a copies nothing(10.3, 123.9)lat, lon = point{1, 2}"grades.csv", "data/grades.csv""r" read, "w" erase and write, "a" add to the end\n that ends each line.strip()continueimport: csvimport csvFinish the Week 3 lab and submit it. The file-reading part is the one to make sure you can do from memory — then repeat the capstone unaided.
Socratica — Python Lists (an excellent mini-lecture) and Corey Schafer — Lists, Tuples & Sets.
Python Tutorial §5 (lists, tuples, sets, dicts) and §7.2 (reading and writing
files); csv module intro — docs.python.org/3.
Everything here is linked on the course page beside this deck.
Next week: organising bigger programs with classes, errors, and project layout.
Classes, errors, and code organisation — how to keep a program readable once it grows past a single notebook cell.
DS 208 · Programming for Data Science