DS 208 · Week 3

Data Structures & Files

Holding many values at once — lists, tuples, sets and dictionaries — then reading and writing real files, up to and including CSV.

Programming for Data Science · University of the Philippines Cebu

Session Map

From one value to many — to files

A

Lists

Order, slices, sorting — and the aliasing trap

B

Tuples & Sets

Fixed records and uniqueness

C

Dictionaries

Lookups, tallies — and nested data

D

Files & CSV

Read, clean, compute, write back

Where We Left Off

Single values only got us so far

Last week: a value, a choice, a function. But real data comes in bulk — thousands of rows. You need containers, and a way to load them.

A container is a value that holds other values — one variable, many items inside.

The four containers

Lists and dictionaries do most of the work; tuples and sets fill two specific gaps.

The missing piece

Getting data in — from a file — so you stop typing it into cells.

Words for Today

Ten words for holding many values — and files

Last week each variable held one value: a number, a string, a boolean. Today’s words are about containers (values that hold other values) and about reading data from files.

list

An ordered collection of values in square brackets. You can change it later.

e.g. [3, 5, 4]

index

An item’s position number, counted from 0 — like seat numbers that start at zero.

e.g. in ["a", "b"], "b" is at index 1

slice

A run of items taken by position: from the start index up to, not including, the stop.

e.g. [10, 20, 30][0:2] is [10, 20]

method

A function that belongs to a value, called with a dot after the value.

e.g. names.append("Ana")

tuple

An ordered collection in round brackets that cannot be changed once made.

e.g. a coordinate (10.3, 123.9)

set

An unordered collection in curly brackets that keeps only one of each value.

e.g. set(["yes", "no", "yes"]) is {'yes', 'no'}

dictionary (dict)

A collection of key → value pairs, like a phone book: look up a name, get a number.

e.g. {"Cebu": 964000, "Davao": 1777000}

key

The label you look a value up by in a dictionary. Each key appears once.

e.g. pop["Cebu"] uses the key "Cebu"

file path

The address of a file on disk. A bare name means “in the current folder”.

e.g. "grades.csv" or "data/grades.csv"

CSV

A plain-text table: one row per line, commas between the columns, often a header row first.

e.g. name,score then Ana,88

Part A

Lists

An ordered shelf you reach by position.

You met lists briefly in Week 2’s loops. This part shows how to reach into one by position, cut pieces out of it, sort it — and the one trap that surprises everybody.

Reach By Position

Counting starts at 0

A list is an ordered collection of values in square brackets. Each item has an index: its position number. The first item is [0], not [1]. Negative indexes count from the end, so [-1] is always the last item.

cities = ["Cebu", "Davao", "Iloilo", "Baguio"] cities[0] # "Cebu" cities[-1] # "Baguio" len(cities) # 4
  • cities = [...]a list of four strings, stored under the name cities
  • cities[0]square brackets after a list mean “the item at this index”: "Cebu"
  • cities[-1]count from the end: the last item, "Baguio"
  • len(cities)how many items: 4. So the last index is 4 - 1 = 3
The Mental Model

Two ways to count the same shelf

"Cebu" "Davao" "Iloilo" "Baguio" 01 23 -4-3 -2-1 index from the start → ← index from the end
The Everyday Moves

Grow, shrink, and check a list

DoCodeResult on [3, 5]
Add to the endnums.append(9)[3, 5, 9]
How manylen(nums)2
Is it in there?5 in numsTrue
Sum themsum(nums)8
Walk each onefor n in nums:3, then 5
The Trap Nobody Warns You About

Two names can share one list

b = a does not copy the list — it gives the same shelf a second name, an alias. Change it through either name and "both" change, because there is only one.

b = a is like giving a friend a second key to your house, not building them a new house. If they move the furniture, your house changes too. .copy() builds the second house.

a = [3, 5, 9] b = a # same list, new name b.append(99) print(a) [3, 5, 9, 99] # a changed too! c = a.copy() # an actual second list c.append(7) # a is untouched

The rule

Want an independent copy? Say so: .copy() (or list(a)).

Putting Things In Order

sorted() returns; .sort() rearranges

Two tools, one crucial difference: sorted(x) hands back a new sorted list; x.sort() reorders the original in place (it changes the list itself rather than making a new one) and returns None.

temps = [31.2, 28.5, 33.1, 29.9] sorted(temps) [28.5, 29.9, 31.2, 33.1] sorted(temps, reverse=True)[:2] [33.1, 31.2] # top two — a classic move print(temps) [31.2, 28.5, 33.1, 29.9] # unchanged

Common bug

x = x.sort() sets x to None. Use one tool or the other, never both at once.

  • sorted(temps)a new list, smallest to largest; temps itself is untouched
  • reverse=Truea keyword argument: sort largest first instead
  • [:2]slice with no start = “from the beginning”: the first two items, i.e. the top two
Strings ↔ Lists

split cuts; join glues

A line of text becomes a list of pieces with .split(); a list becomes one string with .join(). This pair is the bridge between files and containers — remember it for Part D.

line = "Cebu,964000,Visayas" parts = line.split(",") ['Cebu', '964000', 'Visayas'] parts[1] '964000' # a string — not a number yet! " | ".join(["a", "b", "c"]) 'a | b | c'

Note the direction

split is called on the string; join on the separator (the text placed between pieces, here "," or " | "). Everyone trips on that once. The quotes in '964000' say it is still a string.

Quick Check

Tap to reveal

Given nums = [10, 20, 30, 40], what does nums[1:3] return?

A · [10, 20, 30]
B · [20, 30]
C · [20, 30, 40]
D · [10, 20]

B — [20, 30].

Start at index 1 (20), stop before index 3 (40). The stop index is never included — the slice has 3 − 1 = 2 items.

Part B

Tuples & Sets

Two smaller containers that each do one thing perfectly.

You now know lists: ordered and changeable. Each container in this part gives up one list feature on purpose — a tuple cannot change; a set has no order and no duplicates.

Definition · The Frozen List

A tuple is a list that can't change

A tuple uses round brackets instead of square. Index and slice exactly like a list — but it is immutable (cannot be changed after it is made), so any attempt to modify it fails loudly. That rigidity is the feature.

Reading this error

Last line
TypeError: this type (a tuple) does not allow that operation
The message
“item assignment” = changing one item with point[0] = …
Check first
do you need a list instead, or a new tuple?
point = (10.3157, 123.8854) # Cebu City point[0] 10.3157 point[0] = 5 TypeError: 'tuple' object does not support item assignment

When to reach for it

Fixed-shape records: a coordinate, an (R, G, B) colour, a (name, score) pair. Position means something.

The Tuple Superpower

Unpacking: one line, many names

Unpacking: assign a tuple to comma-separated names and Python hands out the values in order — the first to lat, the second to lon. You will use this all through Part C: a dictionary’s .items() hands you (key, value) tuples. (tally below is the vote-count dictionary built in Part C.)

lat, lon = point print(lat) 10.3157 # you've been unpacking all along: for k, v in tally.items(): print(k, v) yes 3

Count must match

Two names need exactly two values — otherwise ValueError: too many/not enough values to unpack. A ValueError means the type was right but the value did not fit (here: the wrong number of items).

Definition · The Uniqueness Machine

A set keeps one of each — unordered

A set: curly brackets, no duplicates, no positions. Two jobs it does better than anything else: dedup (removing duplicates) and fast membership tests (answering “is this value in there?”).

votes = ["yes", "no", "yes", "yes", "no"] set(votes) {'no', 'yes'} # order not guaranteed len(set(votes)) 2 # distinct answers "maybe" in set(votes) False
  • set(votes)make a set from the list: each answer kept once. No index, so the order may vary
  • len(set(votes))how many different answers: 2
  • "maybe" in set(votes)is "maybe" one of them? False

The census question

"How many distinct cities are in this column?" — len(set(column)). One line.

Comparing Groups

Set algebra answers real questions

QuestionOperator{1,2,3} vs {3,4}
In either group (union)A | B{1, 2, 3, 4}
In both groups (intersection)A & B{3}
In A but not B (difference)A - B{1, 2}
The Complete Toolbox

Four containers, four jobs

ContainerLooks likeOrdered?Changeable?Reach for it when…
list[1, 2, 2]YesYesthe items, in order
tuple(10.3, 123.9)YesNoa fixed-shape record
set{1, 2}NoYesuniqueness & membership
dict{"a": 1}InsertionYeslook things up by name
Quick Check

Tap to reveal

A column of 500 barangay names has duplicates. Which expression counts the distinct barangays?

A · len(names)
B · len(set(names))
C · set(len(names))
D · sorted(names)[0]

B — len(set(names)).

set() collapses duplicates; len() counts what's left. A gives 500 (rows, not barangays); C is backwards and errors; D is just the alphabetically first name.

Break

  Ten minutes

Back for dictionaries, nested data, and real files.

Part C

Dictionaries

A labelled drawer you reach by key, not by position.

Lists and tuples find things by position. This part adds the dictionary, which finds things by name — and then puts containers inside containers, which is how real datasets look.

Reach By Key

A name is a better handle than a number

When "the third item" is meaningless but "Davao's population" is exactly what you want, use a dictionary. Each entry is a key : value pair: the key is the label you look up, the value is what you get back.

A dictionary is a phone book: you look up a name (the key) and get a number (the value). Nobody asks for “the 3rd entry in the phone book”.

pop = { "Cebu": 964_000, "Davao": 1_777_000, } pop["Davao"] # 1777000 pop["Baguio"] = 366_000 # add
  • pop = { … }curly brackets with key: value pairs make a dictionary
  • 964_000the underscores are only for your eyes: Python reads 964000
  • pop["Davao"]look up the value stored under the key "Davao": 1777000
  • pop["Baguio"] = 366_000a new key: adds the pair. (An existing key would be overwritten.)
Choosing Between Them

Position or label?

ListDictionary
Reach an item byposition: x[0]key: x["Cebu"]
Order matters?Yes — it's a sequenceInsertion order kept, rarely used
Best when"the items, in order""look this up by name"
Missing itemIndexErrorKeyError (or .get)
The Missing Key

[] crashes; .get() copes

Asking for a key that isn't there raises KeyError. When a gap is expected — and in real data it always is — .get() hands back a default instead.

pop["Manila"] KeyError: 'Manila' pop.get("Manila") # None pop.get("Manila", 0) # 0

Reading this error

Last line
KeyError: 'Manila' — the dictionary has no key 'Manila'; the message is the missing key
Check first
spelling and capitals: "manila" and "Manila" are different keys
.get(key, default)
a method that returns the default (or None) instead of stopping
The Move You'll Use Forever

Tally how often each thing appears

A dictionary maps each distinct item to a running count. This one pattern underlies word counts, vote counts, category frequencies — everything.

votes = ["yes", "no", "yes", "yes"] tally = {} for v in votes: tally[v] = tally.get(v, 0) + 1 tally # {'yes': 3, 'no': 1}
  • tally = {}start with an empty dictionary
  • for v in votes:one pass per vote; v is "yes", then "no"…
  • tally.get(v, 0) + 1the count so far for this vote (0 if it is new), plus one
  • tally[v] = …store the new count under that key
  • {'yes': 3, 'no': 1}the result: three yes votes, one no
Walking A Dictionary

Keys, values, or both

LoopGives youUse when
for k in d:each keyyou only need the labels
for v in d.values():each valueyou only need the numbers
for k, v in d.items():both at onceyou need the pair — the common case
Quick Check

Tap to reveal

You loop for x in scores: over a dictionary scores. What does x hold each time?

A · The values
B · The keys
C · (key, value) pairs
D · An error

B — the keys.

Looping a dict directly gives its keys. For the values use .values(); for both use .items(). Assuming you get values is a very common bug.

Nesting · Where Real Data Lives

A dataset is containers inside containers

Nesting · Navigation

Chain the brackets, outside → in

Each bracket peels one layer. students[0] is a dict; ["name"] reaches into it. Read chains left to right, one hop at a time.

Like a postal address: first the building (students[0]), then the room inside it (["name"]).

students[0] {'name': 'Ana', 'section': 'A', 'score': 88} students[0]["name"] 'Ana' scores["A"][1] 91 # list scores["A"], then index 1

Debug tactic

Lost in a nest? Print one layer at a time: print(x), then print(x[0]), then print(x[0]["name"]).

Nesting · The Working Loop

Filter and find-the-best, by hand

Loop over the rows; each pass holds one dict. Every filter, tally, and "who's top?" question is this loop wearing different clothes. (The grey lines show what the variables hold afterwards.)

  • for s in students:s is one student’s dict each time round
  • s["score"] >= 80that student’s score — at least 80?
  • top = students[0]start by assuming the first student is the best
  • top = sfound a higher score: remember this student instead
# who passed? passing = [] for s in students: if s["score"] >= 80: passing.append(s["name"]) ['Ana', 'Carla'] # who's top? top = students[0] for s in students: if s["score"] > top["score"]: top = s Carla 91
Nesting · The Grouping Move

Build a dict of lists as you loop

The tally pattern's big sibling: instead of counting per key, you collect per key. First time you see a key, start its list; then append.

by_section = {} for s in students: sec = s["section"] if sec not in by_section: by_section[sec] = [] # start the list by_section[sec].append(s["score"]) by_section {'A': [88, 91], 'B': [74]}

This is group-by

When pandas gives you df.groupby("section"), this loop is what it's doing underneath.

  • if sec not in by_section:first time we meet this section? (in on a dict checks its keys)
  • by_section[sec] = []then give it an empty list to fill
  • by_section[sec].append(…)add this student’s score to their section’s list
Part D

Files & CSV

Reading data from disk instead of typing it into a cell.

So far every value was typed into the code. You now have containers to hold data; this part fills them from files on disk, and writes results back out.

Open, Read, Close

A file is read line by line

A file is named data saved on disk; its file path is its address. "hours.txt" alone means “in the current folder”; "data/hours.txt" means “inside the data folder”. Loop over an open file and you get one line per pass. The with block closes the file for you when it ends — even if something errors.

with open("hours.txt") as f: for line in f: print(repr(line)) # '12\n' # '7\n' # '0\n'
  • open("hours.txt")open the file for reading; a wrong path gives FileNotFoundError
  • with … as f:call the open file f inside the indented block; close it at the end
  • for line in f:one line of text per pass — always a string
  • repr(line)show the string with its hidden characters: \n is the end-of-line mark
The Hidden Character

Every line ends in \n

Each line carries an invisible newline character, written \n: the mark that says “the line ends here”. repr() reveals it. Leave it in and comparisons and conversions misbehave in ways you won't see. .strip() removes whitespace (spaces, tabs, newlines) from both ends.

line = "12\n" int(line) # 12 (int is forgiving) line == "12" # False! (the \n) line.strip() # "12" — clean

The habit

.strip() every line as you read it. Remove whitespace before you trust the text.

The Whole Point

Read a file, compute a summary

Strip each line, convert it to a number, accumulate. This tiny loop is the skeleton of every data-loading script you'll ever write.

  • total = count = 0start both at 0 in one line
  • if not line: continuean empty string is falsy; continue = skip the rest of this pass, go to the next line
  • total += int(line)turn the text into a number and add it (+= means total = total + …)
  • print(total / count)the average of the non-blank lines
total = count = 0 with open("hours.txt") as f: for line in f: line = line.strip() if not line: continue # skip blanks total += int(line) count += 1 print(total / count)
Cleaning As You Read

Real files are never perfectly clean

The Other Direction

Writing: mode "w" — and its danger

Open with a mode string, which says what you plan to do: "r" reads (the default), "w" erases and rewrites the file; "a" appends to the end. Newlines are yours to add — write() doesn't.

with open("report.txt", "w") as f: f.write("Mean hours: 6.3\n") f.write("Rows used: 3\n") # "a" adds to the end instead: with open("log.txt", "a") as f: f.write("run finished\n")

Warning worth repeating

open("data.txt", "w") empties the file the instant it runs — before any write. Never open your raw data in "w".

Structured Lines

From lines to columns: split returns

Real data files have several values per line, separated by commas: a CSV file (comma-separated values). The first line is usually the header row — the column names. Part A's .split(",") turns each line into a list — a row.

# grades.csv, line by line name,section,score Ana,A,88 Ben,B,74 line = "Ana,A,88" parts = line.strip().split(",") ['Ana', 'A', '88'] score = int(parts[2]) # strings until you convert
Why There's A Module For This

split(",") breaks on real CSVs

The moment a value itself contains a comma, naive splitting shreds the row. CSV’s quoting rule: a value containing a comma is wrapped in double quotes. The csv module — a file of ready-made code that comes with Python, loaded with import csv — knows all the rules. next(...) takes the first row the reader produces.

line = '"Reyes, Jr.",A,90' line.split(",") ['"Reyes', ' Jr."', 'A', '90'] ✗ 4 pieces! import csv next(csv.reader([line])) ['Reyes, Jr.', 'A', '90'] ✓ correct

The rule

split is fine for files you wrote. Anyone else's CSV: use the csv module.

The csv Module · Rows As Lists

csv.reader: each row arrives split

Wrap the open file; loop rows as lists. next(reader) pulls the header row off the top before the loop.

import csv with open("grades.csv") as f: reader = csv.reader(f) header = next(reader) # take the top row for row in reader: print(row) ['Ana', 'A', '88'] ['Ben', 'B', '74'] …
  • reader = csv.reader(f)wrap the open file: each row will come out as a list of strings
  • header = next(reader)take the first row (the column names) off the top
  • for row in reader:the remaining rows, one list per pass
The csv Module · Rows As Dicts

DictReader: rows you read by name

It uses the header row as keys, so each row is a dict — row["score"] instead of row[2]. This is the list-of-dicts shape from Part C, straight off the disk.

with open("grades.csv") as f: rows = list(csv.DictReader(f)) len(rows) 5 rows[0] {'name': 'Ana', 'section': 'A', 'score': '88'}

Still strings!

'88', not 88. CSV knows nothing about types — int(row["score"]) is on you.

  • list(csv.DictReader(f))read every row as a dict (header names as keys) and collect them in a list
  • len(rows) → 5five data rows; the header is not counted
  • rows[0]the first row: a dict, so rows[0]["name"] is 'Ana'
The csv Module · Writing

Write results back as CSV

csv.writer handles the commas and quoting for you — one writerow per line, header first. w.writerow([...]) takes a list and writes it as one comma-separated line. The grey block shows what ends up in the file.

with open("summary.csv", "w", newline="") as f: w = csv.writer(f) w.writerow(["section", "mean_score"]) w.writerow(["A", 82.0]) w.writerow(["B", 78.0]) section,mean_score A,82.0 B,78.0

That newline=""

Required quirk when writing CSV — without it Windows gets blank lines between rows. Just always include it.

Quick Check

Tap to reveal

After rows = list(csv.DictReader(f)), what is rows[0]["score"] + 2?

A · 90
B · '882'
C · A TypeError
D · '88 + 2'

C — TypeError: can't add str and int.

CSV values are always strings — rows[0]["score"] is '88'. You must convert first: int(rows[0]["score"]) + 2 → 90. (B would be the answer for + "2" — string concatenation.)

Capstone · Live

One program, everything tonight

Read a grades CSV → group scores by section → write a summary CSV. Containers, nesting, files — in twelve lines.

Capstone · The Program

Read, group, average, write

DictReader loads rows; the grouping move collects scores per section; csv.writer saves the answer. Spot all three patterns.

  • rows = list(csv.DictReader(f))1 · read: every row as a dict
  • by_sec[sec].append(int(…))2 · group: each score, converted to a number, into its section’s list
  • for sec in sorted(by_sec):3 · write: the sections in order A, B (sorting a dict sorts its keys)
  • round(… , 1)the mean, rounded to 1 decimal place
import csv with open("grades.csv") as f: rows = list(csv.DictReader(f)) by_sec = {} for r in rows: sec = r["section"] if sec not in by_sec: by_sec[sec] = [] by_sec[sec].append(int(r["score"])) with open("summary.csv", "w", newline="") as f: w = csv.writer(f) w.writerow(["section", "mean_score"]) for sec in sorted(by_sec): w.writerow([sec, round(sum(by_sec[sec]) / len(by_sec[sec]), 1)])
Capstone · The Result

Five rows in, two rows out

This ran for real while building the deck. Input: 5 students, 2 sections. Output: one clean summary another program (or Excel) can open.

# grades.csv (input) name,section,score Ana,A,88 Ben,B,74 Carla,A,91 Diego,B,82 Elena,A,67 # summary.csv (written by our program) section,mean_score A,82.0 B,78.0

Check the arithmetic

Section A: (88 + 91 + 67) / 3 = 82.0. Never trust a pipeline you haven't spot-checked once by hand.

This Week's Lab

Lists, dicts, and a file to total

You'll slice and sort a list, unpack a tuple, count distinct values with a set, tally votes and group scores in a dictionary, write a small file and read it back, then read a CSV and average it per section — the full load-and-summarise loop. ~45 minutes.

How it works: replace each ____, press ▶ Run, then Check. Each part starts with a short reminder of the words it uses.

You'll practise

Indexing and slicing, .get() lookups, the tally pattern, and with open(...).

Then, on your machine

Re-run tonight's capstone on your own CSV — change a score, watch the summary move. Stretch: add a count column per section.

Recap

Four containers, one loop, two directions

Glossary

Week 3 words, one line each

Every new word from today. pandas (Weeks 6–7) is built on all of them.

Before Next Week

Practice & reading

Next Week

Modular Programming

Classes, errors, and code organisation — how to keep a program readable once it grows past a single notebook cell.

DS 208 · Programming for Data Science