Doing maths on thousands of numbers at once — in one line, and far faster than any loop you'd write by hand.
Programming for Data Science · University of the Philippines Cebu
The ndarray: many numbers of one type, packed tight.
Whole-array maths and summaries — no loop required.
Slice, and select by condition with a boolean mask.
NumPy is the engine under pandas, scikit-learn (a machine-learning library), and almost every data tool you'll touch after this course.
Replace a slow Python loop with one fast, readable NumPy line.
A Python for loop over a million prices is slow and wordy. NumPy
moves that work into fast C code and expresses it in a single line.
“C code” means code written in the C language and translated to machine instructions ahead of time. You never write it: NumPy’s inner loops are already written in C, and you call them from Python.
List + loop: many lines, and sluggish on real dataset sizes.
Array + one operation: shorter, clearer, and often 50× faster.
NumPy’s container: many values of one type, stored side by side. “ndarray” = n-dimensional array, its full name.
e.g. np.array([1, 2, 3])
One value inside an array — like one cell of a spreadsheet.
e.g. in [10, 20, 30], 20 is the element at position 1
The data type every element shares: whole numbers, decimals, text…
e.g. int64 (whole numbers), float64 (decimals)
One direction you can move in: a row of numbers has 1 dimension, a table has 2 (down and across).
e.g. [[1, 2], [3, 4]] is 2-D
How many elements along each dimension, written as a tuple (rows, columns).
e.g. (2, 3) = 2 rows, 3 columns
A dimension you name by number when summarising: axis=0 goes down the rows, axis=1 across the columns.
e.g. m.sum(axis=0) gives one total per column
One line of maths that applies to every element at once, with no Python loop.
e.g. prices * 1.12 adds VAT to every price
NumPy stretching a smaller array (or single number) to match a bigger one so they can be combined.
e.g. np.array([1, 2, 3]) + 10 gives [11 12 13]
An array of True/False, one per element, used to keep only the True ones.
e.g. temps[temps > 30] keeps the hot days
Many numbers, one type, packed together.
You already keep numbers in a list and loop over them (weeks 2–3). This part adds the array: the same numbers, packed so that maths runs on all of them at once.
A list can mix types and is flexible but slow for maths. An
array (NumPy calls it an ndarray, “n-dimensional array”)
stores one type in a tight block, so operations run in bulk. Each value in it is an
element.
import numpy as npload the NumPy library under its usual nickname np (an alias, week 4)a = np.array(nums)build an array from the list: the same three numbers in a new containerarray([1, 2, 3])how a notebook shows an array; the word array tells you it is not a listA Python list holds pointers (notes saying where each value really lives) to objects scattered across memory. An array holds the numbers themselves, side by side, all the same type.
Because every element is 8 bytes of float, NumPy can hand the whole block to optimised C code — no per-element type check, no pointer chase.
float64 number takes 8 bytesA list is a coat-check rack: each slot holds a ticket, and you walk off to fetch each coat. An array is eggs in a tray: every egg sits right next to the last, so NumPy can grab the whole tray at once.
☐→objeach list slot holds only a pointer; the numbers themselves live elsewhere[ 1.0 | 2.0 | 3.0 ]the array keeps the numbers themselves, side by sidebudgets.nbytesbudgets is the array of 34,079 flood-control budgets built later today; .nbytes = bytes its numbers use: 34,079 × 8 = 272632 (MB = a million bytes)Same job — add up every flood control budget. One version loops in Python, the other calls .sum().
The second is vectorised: one operation on the whole array, with no Python loop.
At 34k rows it is the difference between instant and noticeable. At 34 million it is the difference between a coffee and going home.
for x in budget_list:budget_list holds the same 34,079 budgets as a plain list; the loop visits them one by one0.714 mstime taken, in milliseconds (thousandths of a second)budgets.sum()one call on the array; NumPy runs the loop inside, in its fast C codePHP 1,600,088,782,393both ways give the same total: about 1.6 trillion pesosAn array’s dtype (data type) is the one type every element shares. Lists let you mix. Arrays do not — and when NumPy has to pick a common type, it may not pick the one you wanted.
Integer division, silent truncation (decimals quietly chopped off) and overflow (a number too big for its type wrapping round to a wrong one) all trace back to a dtype you did not choose.
int64whole numbers, each stored in 64 bits (8 bytes)float64decimal numbers (“floats”); one 3.5 turned 1 and 2 into 1.0 and 2.0<U21text (Unicode strings) up to 21 characters: every element is now a string, so + and mean() stop workingdtype is like the number format of an Excel column: one format for the whole column. Type one word into a column of numbers and the whole column becomes text.
Build them from a range, a count of zeros, or evenly spaced points. These four cover most starting cases. A dimension is one direction you can move in: a row of numbers is 1-D; a table with rows and columns is 2-D.
np.zeros(3)three zeros; 0. is short for 0.0, a floatnp.arange(0, 10, 2)start at 0, step by 2, stop before 10 (like range)np.linspace(0, 1, 5)5 evenly spaced values from 0 to 1, both ends includednp.array([[1, 2], [3, 4]])a list of lists becomes a table: each inner list is one rowYou will almost never type array contents by hand. These four cover nearly every case you meet this term.
arange takes a step and may miss the endpoint with floats. linspace takes a count and always includes it. For plotting, use linspace.
np.ones((2,2))the shape goes in as one tuple (2,2), hence the double bracketsnp.random.default_rng(42)a random-number generator; 42 is the seed: the same seed gives the same numbers every runrng.integers(1, 7, size=5)five whole numbers from 1 to 6 (7 is left out), like five dice rollsreshape lays the same elements out in a new shape (rows × columns). The element count must match. Pass -1 for one dimension and NumPy works it out for you.
So the write-through surprise you will meet in Part C applies here too. If you need independence, .copy().
np.arange(12)the 12 numbers 0 to 11 in one rowa.reshape(3, 4)the same 12 numbers as 3 rows × 4 columns, filled row by rowa.reshape(3, -1)-1 = “work it out”: 12 ÷ 3 = 4 columnsValueError: cannot reshape…read the last line: 5 × 2 = 10 places, but there are 12 elements.shape and .dtype answer the two key questionsHow big is it, and what's inside? The shape is the number of elements along each dimension. An array is one type throughout — mixing in a string quietly turns the whole thing to text.
m = np.array([[1, 2, 3], [4, 5, 6]])two inner lists = two rows of three numbersm.shape(2, 3): a tuple, 2 along the first dimension (rows), 3 along the second (columns)m.dtypeint64: every element is a whole numberm.size6: total elements, 2 × 3Writing a * 2 tells NumPy to double every element in optimised C —
no Python loop runs at all. This is vectorization (also spelled vectorisation).
A loop is a cashier scanning items one by one. A vectorised operation is the whole basket weighed at once: one instruction, every element.
The loop happens once, in C, over tightly packed memory — not in slow Python.
Maths and summaries across the whole array.
You now know what an array is and how to read its shape and dtype. This part does maths with it: one operator or one method call works on every element.
Elementwise means “to each element separately”.
Add, multiply, or take a square root of an array and NumPy does it to each
element and hands back a new array. No loop, no .append().
a * 2every element doubled: [ 2 8 18] (the extra spaces just line numbers up)a + atwo arrays of the same shape are added position by positionnp.sqrt(a)square root of each element; 1. means 1.0, because roots are floatsSum, mean, max, standard deviation — each reduces the whole array at once. These summaries are called aggregations: many values in, one value out. These are the summaries you'll report on every dataset.
temps.sum()all four temperatures added: 121temps.mean()the average: 121 ÷ 4 = 30.25temps.max()the largest value: 33temps.std()standard deviation, the typical distance from the mean: about 1.92 degreesaxis picks the direction to collapse| Call | Collapses | On a 2×3 array gives |
|---|---|---|
m.sum() | everything | one number |
m.sum(axis=0) | down the rows | one per column (3 values) |
m.sum(axis=1) | across the columns | one per row (2 values) |
An axis is a dimension named by number: axis 0 = rows (down), axis 1 = columns (across).
Read it as "the axis that disappears." axis=0 removes the row
dimension, leaving a per-column result.
In a spreadsheet, axis=0 is the “Total” row you add under the columns; axis=1 is the “Total” column you add to the right of the rows.
axis=0 = down. axis=1 = across. The same idea returns in pandas.
Broadcasting lets a single number (a scalar) act on every element without you writing it out. Adding 12% VAT to a whole price list is one multiply.
prices * 1.12the one number 1.12 is stretched to [1.12, 1.12, 1.12]: [112. 224. 336.]prices - prices.mean()the mean (200) is subtracted from every price: [-100. 0. 100.] — each price’s distance from averageBroadcasting is a teacher saying “everyone add 5 marks”: one instruction, copied out to every student, without writing 5 next to each name.
Given a = np.array([1, 2, 3]), what
is a * 3?
[1, 2, 3, 1, 2, 3, 1, 2, 3][3, 6, 9]6B — [3, 6, 9].
On an array, * 3 multiplies every
element. On a Python list, * 3 would repeat it (answer A) — a
key difference between the two.
Array maths is elementwise; list * is repetition. Know which
object you hold.
Array * n scales; list * n repeats.
Back to pull out exactly the values you want.
Slice by position, or select by condition.
You already slice lists with nums[1:3] (week 3). This part uses the same brackets on arrays — one index per dimension — and adds choosing values by a condition.
One index per axis, separated by a comma. A lone : means "all of
this axis" — so m[:, 0] is the whole first column.
a[1:3]positions 1 and 2 (a slice stops before the end number): [20 30]m[0, 1]row 0, column 1 — counting from 0, like list positions: 2m[:, 0]: = every row; 0 = column 0: [1 3]A comparison returns an array of True/False (booleans) — a
boolean mask. Index with it and you
keep only the elements where it's True. This is filtering.
temps > 30compare every element with 30: [F T T F] is short for [False True True False]temps[temps > 30]use that mask inside the brackets: keep the elements where it is True, [31 33]A mask is a stencil laid over the array: the holes (True) let values through, the solid parts (False) block them.
It is a window onto the same memory. Writing through it changes the original.
Slicing an array does not copy the data. You get a view (a window onto the same numbers) — and assigning into it writes straight through to the array you sliced.
Copying a large array is expensive. A view costs nothing, which is what makes chained slicing cheap. The price is this surprise.
s = a[1:4]s is a view: a window onto positions 1–3 of as[0] = 999write into the window…print(a)…and a[1] changed too: [ 1 999 3 4 5]s.base is a.base = the array a view looks through; True means s is a view of aFancy indexing means indexing with a list of positions instead of a slice. NumPy has to build a new array. Now the original is safe.
Slice a[1:4] → view. List a[[1,2,3]] or mask a[a>2] → copy. When in doubt, .copy() explicitly.
f = a[[1,2,3]]outer brackets index, inner brackets are a list of positions: a new array (a copy)print(a)changing f left a alone: [1 2 3 4 5]f.base is NoneNone = “no base”: f owns its own numbersThe memory hook: axis=0 goes down, axis=1 goes across — and the axis you name is the one that disappears from the shape.
NumPy compares shapes from the right. Each pair must be equal, or one of them must be 1. The “trailing” size is the last number in each shape.
Nearly every shape bug is solved by printing .shape on both sides before the operation.
np.ones((3,4)) + np.array([1,2,3,4])shapes (3, 4) and (4,): last sizes 4 and 4 match, so the row of 4 is added to all 3 rowsnp.ones((3,4)) + np.array([1,2,3])shapes (3, 4) and (3,): last sizes 4 and 3 differ and neither is 1ValueError: … (3,4) (3,)read the last line: the error type, then the two shapes that could not be lined up (“operands” = the two arrays on either side of +)Three arguments: a condition, a value where it is true, a value where it is false. It returns a new array.
939 flood projects list a budget of 0, which really means “not recorded”. np.where(b == 0, np.nan, b) marks them all as missing in one line.
x < 0the condition, checked for every element: [F T F T]np.where(x < 0, 0, x)where it is True use 0, otherwise keep x: [10 0 3 0]An array is a fixed block of memory. There is no room at the end, so every np.append allocates (reserves) a whole new array and copies everything across.
Build with a Python list and convert once, or allocate with np.zeros(n) and fill by index. Never grow an array in a loop.
out = np.append(out, i)each pass builds a brand-new array one element longer: 10,000 copiesnp.array([i for i in range(10000)])a comprehension (the one-line loop from week 4) grows a list cheaply; then convert to an array oncenp.arange(10000)best of all: NumPy makes the numbers directlycould not be broadcast togetherYour trailing dimensions do not match and neither is 1. Print both .shapes — the answer is always in that pair.
cannot reshape array of size 12The element count changed. reshape never invents or discards elements; check whether you meant to filter first.
too many indices for arrayYou used a[i, j] on a 1-D array. It has one axis, so it takes one index.
NumPy’s errors are unusually good: they name the shapes involved. The habit worth building is reading the last line before you change anything.
ValueError (impossible shape); the third is an IndexError (a position that does not exist).shape of each array used on itMake a = np.arange(10). Slice b = a[2:5]. Set b[0] = -1. Print a. Explain what happened.
Redo it so a survives. Two ways — find both.
Build a 3×4 array. Produce one total per column, then one per row. Check the shapes match what you expected.
Add a shape that cannot broadcast. Read the error aloud — it names both shapes.
If step 1 does not change a, you used fancy indexing by accident. That is a useful mistake to have made once.
Filter to the wet days, average them, and count the dry ones — three questions, three one-liners, no loop anywhere.
wet = rain[rain > 0]mask: keep the days with some rain, [12 4 33 8]wet.mean()(12 + 4 + 33 + 8) ÷ 4 = 14.25 mm on an average wet day(rain == 0).sum()True counts as 1 and False as 0, so summing the mask counts the dry days: 2The same array operations, on a real public dataset: the government’s (DPWH) list of flood control projects, which DS 227, the companion course, is cleaning this week.
Every summary on the next slides is one call on this array. No loops appear anywhere.
Always. They tell you whether the load worked before you trust any number that follows.
float(r["budget"]) for r in rowsrows is the CSV read with csv.DictReader (week 3); take each row’s budget text and make it a float((34079,), dtype('float64'))1-D, 34,079 elements, all decimals — the load worked1600088782392.8298the grand total in pesos: about 1.6 trillion (1,600,088,782,392.83)That last trick is worth keeping: a comparison gives an array of True/False, and summing it counts the trues.
A boolean array used as an index keeps exactly the elements where it is True. This is the move you will use most.
Combine with & and | — and keep the parentheses, because & binds tighter than >.
np.sort(big)[::-1]sort smallest-first, then [::-1] reverses it: biggest first1.44750000e+09scientific notation: 1.4475 × 109, about 1.45 billion pesos(budgets > 1e8) & (budgets < 5e8)& = “and” for masks: between 100 million and 500 millionmid.sizehow many elements passed both tests: 2145 projectsargsort returns the positions that would put the array in order. When you need "the top ten projects", not "the top ten numbers", you need the indices — so you can look up the other columns.
order = arr.argsort()[::-1] then [rows[i] for i in order[:10]]. Sort once, index anything.
budgets.argsort()[::-1]positions from smallest to biggest, reversed: biggest firstorder[:3]the first three positions: where the three biggest budgets sitrows[i]["region"]use each position to look up that project’s other columns in rows| Slice | Projects | Share of ₱1.6T |
|---|---|---|
| Top 1% | 341 | 6.5% |
| Over ₱1 billion | 3 | 0.2% |
| Below the median | 17,039 | 16.5% |
Three lines of NumPy answered a question about public spending concentration. That is the whole point of the tool — the analysis is the easy part once the data is an array.
np.sort(budgets)[::-1][:341].sum() / budgets.sum() — sort biggest-first, keep the first 341 (1% of 34,079), add them up, divide by the grand total.
You slice b = a[0:3] and then run b[:] = 0. What happens to a?
b is a separate copyb changes, unless you call .flush()a. Check with b.base is a → True. Use a[0:3].copy() when you want independence.m has shape (2, 3). What shape does m.sum(axis=0) return?
(2,) — one total per row(3,) — one total per column(2, 3) — unchanged(3,). axis=0 goes down, axis=1 goes across.Bulk operations in C — often tens of times faster than a Python loop.
rain[rain > 0] reads like the question you're asking.
pandas columns are arrays; scikit-learn (a machine-learning library) eats arrays. Learn it once, reuse it always.
Next week's pandas Series is a labelled NumPy array — the maths
you just learned carries straight over.
Master masks now; you'll filter DataFrames the exact same way.
For a = np.array([5, 12, 7, 20]),
what does a[a > 10] return?
[True, True][12, 20][5, 7]2B — [12, 20].
a > 10 is the mask
[F T F T]; indexing with it keeps the elements where it's
True. To count them instead, use
(a > 10).sum().
Mask to keep values; .sum() the mask to count
them.
a[mask] selects; mask.sum() counts.
You'll build arrays, do elementwise maths, take axis-wise summaries of a 2-D array, and filter real weather data with a boolean mask. ~45 minutes.
np.array, arange, aggregations, axis=, and
a[a > k].
Compare a loop's runtime against the vectorized version with %timeit (a notebook command that times one line).
One type, packed tight. Check .shape and .dtype.
Elementwise maths and axis-wise summaries — no loops.
Slice by position; select by condition with a boolean mask.
One sentence: put your numbers in an array and describe what you want — NumPy does the loop for you, fast.
Compute over whole datasets in single, readable lines.
int64, float64, <U21 (text)(2, 3) = 2 rows, 3 columnssum, mean, max, std0 down the rows, 1 acrossa[mask] keeps the True onesFinish the Week 5 lab and submit it. The boolean-mask section is the one to be able to do from memory.
NumPy "Absolute Beginners" guide — numpy.org/doc/stable/user/absolute_beginners.
Everything here is linked on the course page beside this deck.
Next week: pandas gives these arrays column names and an index.
The DataFrame — a spreadsheet in code. Loading a CSV, then selecting and filtering the rows and columns you actually need.
DS 208 · Programming for Data Science