DS 208 · Week 5

Scientific Computing with NumPy

Doing maths on thousands of numbers at once — in one line, and far faster than any loop you'd write by hand.

Programming for Data Science · University of the Philippines Cebu

Session Map

One new object, three superpowers

Where We Left Off

Loops work — until the data gets big

A Python for loop over a million prices is slow and wordy. NumPy moves that work into fast C code and expresses it in a single line.

“C code” means code written in the C language and translated to machine instructions ahead of time. You never write it: NumPy’s inner loops are already written in C, and you call them from Python.

The pain

List + loop: many lines, and sluggish on real dataset sizes.

The gain

Array + one operation: shorter, clearer, and often 50× faster.

Words for Today

Nine words you will hear this week

array (ndarray)

NumPy’s container: many values of one type, stored side by side. “ndarray” = n-dimensional array, its full name.

e.g. np.array([1, 2, 3])

element

One value inside an array — like one cell of a spreadsheet.

e.g. in [10, 20, 30], 20 is the element at position 1

dtype

The data type every element shares: whole numbers, decimals, text…

e.g. int64 (whole numbers), float64 (decimals)

dimension

One direction you can move in: a row of numbers has 1 dimension, a table has 2 (down and across).

e.g. [[1, 2], [3, 4]] is 2-D

shape

How many elements along each dimension, written as a tuple (rows, columns).

e.g. (2, 3) = 2 rows, 3 columns

axis

A dimension you name by number when summarising: axis=0 goes down the rows, axis=1 across the columns.

e.g. m.sum(axis=0) gives one total per column

vectorised operation

One line of maths that applies to every element at once, with no Python loop.

e.g. prices * 1.12 adds VAT to every price

broadcasting

NumPy stretching a smaller array (or single number) to match a bigger one so they can be combined.

e.g. np.array([1, 2, 3]) + 10 gives [11 12 13]

boolean mask

An array of True/False, one per element, used to keep only the True ones.

e.g. temps[temps > 30] keeps the hot days

Part A

The Array

Many numbers, one type, packed together.

You already keep numbers in a list and loop over them (weeks 2–3). This part adds the array: the same numbers, packed so that maths runs on all of them at once.

List vs Array

A list holds anything; an array holds numbers well

A list can mix types and is flexible but slow for maths. An array (NumPy calls it an ndarray, “n-dimensional array”) stores one type in a tight block, so operations run in bulk. Each value in it is an element.

import numpy as np nums = [1, 2, 3] # a Python list a = np.array(nums) # a NumPy array a # array([1, 2, 3])
  • import numpy as npload the NumPy library under its usual nickname np (an alias, week 4)
  • a = np.array(nums)build an array from the list: the same three numbers in a new container
  • array([1, 2, 3])how a notebook shows an array; the word array tells you it is not a list
The Reason

Why an array is fast, in one picture

A Python list holds pointers (notes saying where each value really lives) to objects scattered across memory. An array holds the numbers themselves, side by side, all the same type.

One type, known width

Because every element is 8 bytes of float, NumPy can hand the whole block to optimised C code — no per-element type check, no pointer chase.

New words

byte
a small unit of computer memory; one float64 number takes 8 bytes
memory
where a running program keeps its values; “tight block” = one unbroken stretch of it

A list is a coat-check rack: each slot holds a ticket, and you walk off to fetch each coat. An array is eggs in a tray: every egg sits right next to the last, so NumPy can grab the whole tray at once.

# a list: 34,079 pointers to 34,079 objects [1.0, 2.0, 3.0] → ☐→obj ☐→obj ☐→obj # an array: one block of 8-byte floats np.array([1.0, 2.0, 3.0]) → [ 1.0 | 2.0 | 3.0 ] # same numbers, 4x less memory budgets.nbytes 272632 # vs ~1.09 MB as a list
  • ☐→objeach list slot holds only a pointer; the numbers themselves live elsewhere
  • [ 1.0 | 2.0 | 3.0 ]the array keeps the numbers themselves, side by side
  • budgets.nbytesbudgets is the array of 34,079 flood-control budgets built later today; .nbytes = bytes its numbers use: 34,079 × 8 = 272632 (MB = a million bytes)
Measured, Not Claimed

146× on the same 34,079 numbers

Same job — add up every flood control budget. One version loops in Python, the other calls .sum(). The second is vectorised: one operation on the whole array, with no Python loop.

This gap grows with the data

At 34k rows it is the difference between instant and noticeable. At 34 million it is the difference between a coffee and going home.

# the loop you would write first total = 0.0 for x in budget_list: total += x 0.714 ms # the same answer, vectorised budgets.sum() 0.005 ms # 146x faster PHP 1,600,088,782,393
  • for x in budget_list:budget_list holds the same 34,079 budgets as a plain list; the loop visits them one by one
  • 0.714 mstime taken, in milliseconds (thousandths of a second)
  • budgets.sum()one call on the array; NumPy runs the loop inside, in its fast C code
  • PHP 1,600,088,782,393both ways give the same total: about 1.6 trillion pesos
Type Matters

An array has one dtype, and it can bite

An array’s dtype (data type) is the one type every element shares. Lists let you mix. Arrays do not — and when NumPy has to pick a common type, it may not pick the one you wanted.

Check dtype when maths looks wrong

Integer division, silent truncation (decimals quietly chopped off) and overflow (a number too big for its type wrapping round to a wrong one) all trace back to a dtype you did not choose.

np.array([1, 2, 3]).dtype int64 # one float makes the whole array float np.array([1, 2, 3.5]).dtype float64 # one string makes it ALL strings np.array([1, 2, '3']).dtype <U21 # now maths is broken
  • int64whole numbers, each stored in 64 bits (8 bytes)
  • float64decimal numbers (“floats”); one 3.5 turned 1 and 2 into 1.0 and 2.0
  • <U21text (Unicode strings) up to 21 characters: every element is now a string, so + and mean() stop working

dtype is like the number format of an Excel column: one format for the whole column. Type one word into a column of numbers and the whole column becomes text.

Four Ways To Start

You rarely type arrays by hand

Build them from a range, a count of zeros, or evenly spaced points. These four cover most starting cases. A dimension is one direction you can move in: a row of numbers is 1-D; a table with rows and columns is 2-D.

np.zeros(3) # [0. 0. 0.] np.arange(0, 10, 2) # [0 2 4 6 8] np.linspace(0, 1, 5) # [0. .25 .5 .75 1.] np.array([[1, 2], [3, 4]]) # 2-D
  • np.zeros(3)three zeros; 0. is short for 0.0, a float
  • np.arange(0, 10, 2)start at 0, step by 2, stop before 10 (like range)
  • np.linspace(0, 1, 5)5 evenly spaced values from 0 to 1, both ends included
  • np.array([[1, 2], [3, 4]])a list of lists becomes a table: each inner list is one row
Without Typing Them

Four ways to conjure an array

You will almost never type array contents by hand. These four cover nearly every case you meet this term.

linspace vs arange

arange takes a step and may miss the endpoint with floats. linspace takes a count and always includes it. For plotting, use linspace.

np.arange(0, 10, 2) [0 2 4 6 8] np.linspace(0, 1, 5) [0. 0.25 0.5 0.75 1. ] np.zeros(3), np.ones((2,2)) [0. 0. 0.] [[1. 1.] [1. 1.]] rng = np.random.default_rng(42) rng.integers(1, 7, size=5) [1 5 4 3 3] # seeded = repeatable
  • np.ones((2,2))the shape goes in as one tuple (2,2), hence the double brackets
  • np.random.default_rng(42)a random-number generator; 42 is the seed: the same seed gives the same numbers every run
  • rng.integers(1, 7, size=5)five whole numbers from 1 to 6 (7 is left out), like five dice rolls
Same Data, New Shape

reshape rearranges; it does not recompute

reshape lays the same elements out in a new shape (rows × columns). The element count must match. Pass -1 for one dimension and NumPy works it out for you.

reshape usually returns a view

So the write-through surprise you will meet in Part C applies here too. If you need independence, .copy().

New word

view
a new way of looking at the same numbers: change one and the other changes too (Part C shows it)
a = np.arange(12) a.reshape(3, 4) [[ 0 1 2 3] [ 4 5 6 7] [ 8 9 10 11]] a.reshape(3, -1).shape (3, 4) # -1 = "you figure it out" a.reshape(5, 2) ValueError: cannot reshape array of size 12 into shape (5,2)
  • np.arange(12)the 12 numbers 0 to 11 in one row
  • a.reshape(3, 4)the same 12 numbers as 3 rows × 4 columns, filled row by row
  • a.reshape(3, -1)-1 = “work it out”: 12 ÷ 3 = 4 columns
  • ValueError: cannot reshape…read the last line: 5 × 2 = 10 places, but there are 12 elements
Know Your Array

.shape and .dtype answer the two key questions

How big is it, and what's inside? The shape is the number of elements along each dimension. An array is one type throughout — mixing in a string quietly turns the whole thing to text.

m = np.array([[1, 2, 3], [4, 5, 6]]) m.shape # (2, 3) -> 2 rows, 3 cols m.dtype # int64 m.size # 6 total elements
  • m = np.array([[1, 2, 3], [4, 5, 6]])two inner lists = two rows of three numbers
  • m.shape(2, 3): a tuple, 2 along the first dimension (rows), 3 along the second (columns)
  • m.dtypeint64: every element is a whole number
  • m.size6: total elements, 2 × 3
The Big Idea

One operation, applied to everything

Loop: one at a time 1 2 3 … step, step, step — slow Vectorized: all at once 1 2 3 … a * 2 — one instruction
Part B

Operations

Maths and summaries across the whole array.

You now know what an array is and how to read its shape and dtype. This part does maths with it: one operator or one method call works on every element.

Elementwise

Operators apply to every element

Elementwise means “to each element separately”. Add, multiply, or take a square root of an array and NumPy does it to each element and hands back a new array. No loop, no .append().

a = np.array([1, 4, 9]) a * 2 # [ 2 8 18] a + a # [ 2 8 18] np.sqrt(a) # [1. 2. 3.]
  • a * 2every element doubled: [ 2 8 18] (the extra spaces just line numbers up)
  • a + atwo arrays of the same shape are added position by position
  • np.sqrt(a)square root of each element; 1. means 1.0, because roots are floats
Summaries In One Call

Collapse an array to a single number

Sum, mean, max, standard deviation — each reduces the whole array at once. These summaries are called aggregations: many values in, one value out. These are the summaries you'll report on every dataset.

temps = np.array([28, 31, 33, 29]) temps.sum() # 121 temps.mean() # 30.25 temps.max() # 33 temps.std() # 1.92
  • temps.sum()all four temperatures added: 121
  • temps.mean()the average: 121 ÷ 4 = 30.25
  • temps.max()the largest value: 33
  • temps.std()standard deviation, the typical distance from the mean: about 1.92 degrees
2-D Summaries

axis picks the direction to collapse

CallCollapsesOn a 2×3 array gives
m.sum()everythingone number
m.sum(axis=0)down the rowsone per column (3 values)
m.sum(axis=1)across the columnsone per row (2 values)
Mixing Shapes

A scalar stretches to fit the array

Broadcasting lets a single number (a scalar) act on every element without you writing it out. Adding 12% VAT to a whole price list is one multiply.

prices = np.array([100, 200, 300]) prices * 1.12 # [112. 224. 336.] prices - prices.mean() # centred
  • prices * 1.12the one number 1.12 is stretched to [1.12, 1.12, 1.12]: [112. 224. 336.]
  • prices - prices.mean()the mean (200) is subtracted from every price: [-100. 0. 100.] — each price’s distance from average

Broadcasting is a teacher saying “everyone add 5 marks”: one instruction, copied out to every student, without writing 5 next to each name.

Quick Check

Tap to reveal

Given a = np.array([1, 2, 3]), what is a * 3?

A · [1, 2, 3, 1, 2, 3, 1, 2, 3]
B · [3, 6, 9]
C · 6
D · an error

B — [3, 6, 9].

On an array, * 3 multiplies every element. On a Python list, * 3 would repeat it (answer A) — a key difference between the two.

Break

  Five minutes

Back to pull out exactly the values you want.

Part C

Indexing & Selecting

Slice by position, or select by condition.

You already slice lists with nums[1:3] (week 3). This part uses the same brackets on arrays — one index per dimension — and adds choosing values by a condition.

By Position

Slicing works in each dimension

One index per axis, separated by a comma. A lone : means "all of this axis" — so m[:, 0] is the whole first column.

a = np.array([10, 20, 30, 40]) a[1:3] # [20 30] m = np.array([[1, 2], [3, 4]]) m[0, 1] # 2 (row 0, col 1) m[:, 0] # [1 3] first column
  • a[1:3]positions 1 and 2 (a slice stops before the end number): [20 30]
  • m[0, 1]row 0, column 1 — counting from 0, like list positions: 2
  • m[:, 0]: = every row; 0 = column 0: [1 3]
By Condition

The boolean mask — NumPy's best trick

A comparison returns an array of True/False (booleans) — a boolean mask. Index with it and you keep only the elements where it's True. This is filtering.

temps = np.array([28, 31, 33, 29]) temps > 30 # [F T T F] temps[temps > 30] # [31 33]
  • temps > 30compare every element with 30: [F T T F] is short for [False True True False]
  • temps[temps > 30]use that mask inside the brackets: keep the elements where it is True, [31 33]

A mask is a stencil laid over the array: the holes (True) let values through, the solid parts (False) block them.

The Trap Nobody Warns You About

A slice is not a copy

It is a window onto the same memory. Writing through it changes the original.

Views

a[1:4] hands you a window, not a photograph

Slicing an array does not copy the data. You get a view (a window onto the same numbers) — and assigning into it writes straight through to the array you sliced.

Why NumPy does this

Copying a large array is expensive. A view costs nothing, which is what makes chained slicing cheap. The price is this surprise.

a = np.array([1,2,3,4,5]) s = a[1:4] # a VIEW s[0] = 999 print(a) [ 1 999 3 4 5] # you never touched `a` by name s.base is a True # the giveaway
  • s = a[1:4]s is a view: a window onto positions 1–3 of a
  • s[0] = 999write into the window…
  • print(a)…and a[1] changed too: [ 1 999 3 4 5]
  • s.base is a.base = the array a view looks through; True means s is a view of a
Copies

Fancy indexing does copy — so it behaves differently

Fancy indexing means indexing with a list of positions instead of a slice. NumPy has to build a new array. Now the original is safe.

The rule

Slice a[1:4] → view. List a[[1,2,3]] or mask a[a>2] → copy. When in doubt, .copy() explicitly.

a = np.array([1,2,3,4,5]) f = a[[1,2,3]] # a COPY f[0] = 999 print(a) [1 2 3 4 5] # unchanged f.base is None True
  • f = a[[1,2,3]]outer brackets index, inner brackets are a list of positions: a new array (a copy)
  • print(a)changing f left a alone: [1 2 3 4 5]
  • f.base is NoneNone = “no base”: f owns its own numbers
The Axis Argument

axis says what to collapse, not what to keep

m = np.array([[1,2,3], [4,5,6]]) m.shape (2, 3) m.sum() 21 # everything
m.sum(axis=0) [5 7 9] # shape (3,) # collapsed the 2 ROWS # one total per column m.sum(axis=1) [ 6 15] # shape (2,) # collapsed the 3 COLUMNS # one total per row
Broadcasting

A smaller shape stretches — if the trailing sizes line up

NumPy compares shapes from the right. Each pair must be equal, or one of them must be 1. The “trailing” size is the last number in each shape.

Read the error, it tells you the shapes

Nearly every shape bug is solved by printing .shape on both sides before the operation.

# works: trailing 4 matches 4 np.ones((3,4)) + np.array([1,2,3,4]) -> shape (3, 4) # fails: trailing 4 vs 3 np.ones((3,4)) + np.array([1,2,3]) ValueError: operands could not be broadcast together with shapes (3,4) (3,)
  • np.ones((3,4)) + np.array([1,2,3,4])shapes (3, 4) and (4,): last sizes 4 and 4 match, so the row of 4 is added to all 3 rows
  • np.ones((3,4)) + np.array([1,2,3])shapes (3, 4) and (3,): last sizes 4 and 3 differ and neither is 1
  • ValueError: … (3,4) (3,)read the last line: the error type, then the two shapes that could not be lined up (“operands” = the two arrays on either side of +)
Vectorised if/else

np.where replaces the loop you were about to write

Three arguments: a condition, a value where it is true, a value where it is false. It returns a new array.

Use it to clean, too

939 flood projects list a budget of 0, which really means “not recorded”. np.where(b == 0, np.nan, b) marks them all as missing in one line.

New word

NaN / np.nan
“Not a Number”: NumPy’s marker for a missing value. Sums and means that include it give NaN, so it cannot hide
x = np.array([10, -5, 3, -8]) np.where(x < 0, 0, x) [10 0 3 0] # condition, if-true, if-false # no loop, no append, no index
  • x < 0the condition, checked for every element: [F T F T]
  • np.where(x < 0, 0, x)where it is True use 0, otherwise keep x: [10 0 3 0]
The Performance Trap

np.append in a loop is slower than the list you replaced

An array is a fixed block of memory. There is no room at the end, so every np.append allocates (reserves) a whole new array and copies everything across.

Reading the comments

O(n^2)
shorthand for “work grows with the square of the size”: 10× more data, about 100× more time
O(n)
work grows in step with the size: 10× more data, 10× more time

The fix

Build with a Python list and convert once, or allocate with np.zeros(n) and fill by index. Never grow an array in a loop.

# O(n^2) — 10,000 copies of a growing array out = np.array([]) for i in range(10000): out = np.append(out, i) 14.7 ms # O(n) — build a list, convert once out = np.array([i for i in range(10000)]) 0.257 ms # 57x faster # and if a range is all you need out = np.arange(10000) 0.001 ms
  • out = np.append(out, i)each pass builds a brand-new array one element longer: 10,000 copies
  • np.array([i for i in range(10000)])a comprehension (the one-line loop from week 4) grows a list cheaply; then convert to an array once
  • np.arange(10000)best of all: NumPy makes the numbers directly
When The Shapes Fight

Three errors, and what each one is really telling you

Your Turn · 8 min

Make an array misbehave, then fix it

1 · Prove the view

Make a = np.arange(10). Slice b = a[2:5]. Set b[0] = -1. Print a. Explain what happened.

2 · Defend against it

Redo it so a survives. Two ways — find both.

3 · Get the axis right

Build a 3×4 array. Produce one total per column, then one per row. Check the shapes match what you expected.

4 · Break broadcasting on purpose

Add a shape that cannot broadcast. Read the error aloud — it names both shapes.

Putting It Together

Summarise the days it rained

Filter to the wet days, average them, and count the dry ones — three questions, three one-liners, no loop anywhere.

rain = np.array([0, 12, 4, 33, 0, 8]) wet = rain[rain > 0] wet.mean() # 14.25 mm (rain == 0).sum() # 2 dry days
  • wet = rain[rain > 0]mask: keep the days with some rain, [12 4 33 8]
  • wet.mean()(12 + 4 + 33 + 8) ÷ 4 = 14.25 mm on an average wet day
  • (rain == 0).sum()True counts as 1 and False as 0, so summing the mask counts the dry days: 2
Now On Real Money

34,079 flood control budgets

The same array operations, on a real public dataset: the government’s (DPWH) list of flood control projects, which DS 227, the companion course, is cleaning this week.

One Column, One Array

From CSV to a NumPy array

Every summary on the next slides is one call on this array. No loops appear anywhere.

Shape and dtype first

Always. They tell you whether the load worked before you trust any number that follows.

budgets = np.array([float(r["budget"]) for r in rows]) budgets.shape, budgets.dtype ((34079,), dtype('float64')) budgets.sum() 1600088782392.8298 # PHP 1.6 trillion
  • float(r["budget"]) for r in rowsrows is the CSV read with csv.DictReader (week 3); take each row’s budget text and make it a float
  • ((34079,), dtype('float64'))1-D, 34,079 elements, all decimals — the load worked
  • 1600088782392.8298the grand total in pesos: about 1.6 trillion (1,600,088,782,392.83)
Four Questions, Four Lines

No loop appears anywhere

budgets.mean() 46,952,340 np.median(budgets) 37,676,233 # mean > median => right skew
budgets.max() 1,447,499,996 (budgets > 1e9).sum() 3 # over PHP 1B # a bool array sums as 0/1
The Mask, On Public Money

Select rows by a condition, not a position

A boolean array used as an index keeps exactly the elements where it is True. This is the move you will use most.

Masks compose

Combine with & and | — and keep the parentheses, because & binds tighter than >.

big = budgets[budgets > 1e9] np.sort(big)[::-1] array([1.44750000e+09, 1.09625523e+09, 1.04278280e+09]) # mid-size band: note the () mid = budgets[(budgets > 1e8) & (budgets < 5e8)] mid.size 2145
  • np.sort(big)[::-1]sort smallest-first, then [::-1] reverses it: biggest first
  • 1.44750000e+09scientific notation: 1.4475 × 109, about 1.45 billion pesos
  • (budgets > 1e8) & (budgets < 5e8)& = “and” for masks: between 100 million and 500 million
  • mid.sizehow many elements passed both tests: 2145 projects
Ranking

argsort gives you positions, not values

argsort returns the positions that would put the array in order. When you need "the top ten projects", not "the top ten numbers", you need the indices — so you can look up the other columns.

The pattern to memorise

order = arr.argsort()[::-1] then [rows[i] for i in order[:10]]. Sort once, index anything.

order = budgets.argsort()[::-1] order[:3] array([18777, 33185, 28049]) # POSITIONS of the 3 biggest # now look up the real records for i in order[:3]: print(rows[i]["region"]) Central Office Central Office Central Office
  • budgets.argsort()[::-1]positions from smallest to biggest, reversed: biggest first
  • order[:3]the first three positions: where the three biggest budgets sit
  • rows[i]["region"]use each position to look up that project’s other columns in rows
Concentration

The top 1% of projects hold 6.5% of all spending

SliceProjectsShare of ₱1.6T
Top 1%3416.5%
Over ₱1 billion30.2%
Below the median17,03916.5%
Quick Check

Tap to reveal

You slice b = a[0:3] and then run b[:] = 0. What happens to a?

A · Nothing — b is a separate copy
B · Its first three elements become 0, because a slice is a view
C · It raises a ValueError about read-only arrays
D · Only b changes, unless you call .flush()
B. Basic slicing returns a view that shares memory with the original, so writing through it changes a. Check with b.base is a → True. Use a[0:3].copy() when you want independence.
Quick Check

Tap to reveal

m has shape (2, 3). What shape does m.sum(axis=0) return?

A · (2,) — one total per row
B · (3,) — one total per column
C · (2, 3) — unchanged
D · a single number
B. The axis you name is the one that disappears. Naming axis 0 collapses the 2 rows, leaving one total per column — shape (3,). axis=0 goes down, axis=1 goes across.
Why This Sticks

The foundation everything sits on

Quick Check

Tap to reveal

For a = np.array([5, 12, 7, 20]), what does a[a > 10] return?

A · [True, True]
B · [12, 20]
C · [5, 7]
D · 2

B — [12, 20].

a > 10 is the mask [F T F T]; indexing with it keeps the elements where it's True. To count them instead, use (a > 10).sum().

This Week's Lab

Arrays, summaries, and a mask

You'll build arrays, do elementwise maths, take axis-wise summaries of a 2-D array, and filter real weather data with a boolean mask. ~45 minutes.

You'll practise

np.array, arange, aggregations, axis=, and a[a > k].

Stretch, if you want

Compare a loop's runtime against the vectorized version with %timeit (a notebook command that times one line).

Recap

Array in, one line, answer out

Glossary Recap

Every new word from today, one line each

The array

array (ndarray)
NumPy’s container: many values of one type, side by side
element
one value inside an array
dtype
the one type all elements share: int64, float64, <U21 (text)
dimension
one direction to move in: 1-D row, 2-D table
shape
elements per dimension, e.g. (2, 3) = 2 rows, 3 columns
pointer / byte
a note saying where a value lives / a small unit of memory
C code
fast, pre-translated code NumPy runs for you behind the scenes
reshape
lay the same elements out in a new shape
seed
a starting number that makes “random” numbers repeatable

Maths and selecting

vectorised operation
one line of maths applied to every element, no Python loop
elementwise
done to each element separately, position by position
aggregation
many values in, one out: sum, mean, max, std
axis
which dimension to collapse: 0 down the rows, 1 across
broadcasting
stretching a smaller array or number to fit a bigger one
boolean mask
True/False per element; a[mask] keeps the True ones
view / copy
a window onto the same numbers / an independent duplicate
fancy indexing
indexing with a list of positions; always makes a copy
np.where
vectorised if/else: condition, value-if-true, value-if-false
NaN
“Not a Number”: the marker for a missing value
argsort
the positions that would sort an array
Before Next Week

Practice & reading

Next Week

Data Manipulation with pandas I

The DataFrame — a spreadsheet in code. Loading a CSV, then selecting and filtering the rows and columns you actually need.

DS 208 · Programming for Data Science