DS 227 · Week 7

Data Exploration I

Descriptive statistics and univariate EDA — before any model, you describe: one variable at a time, honestly.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

One variable, three questions

Words for Today · 1 of 2

Nine words for describing one column

EDA

Exploratory data analysis: looking at data with summaries and charts before testing any idea.

e.g. run describe() and draw a histogram before any model

variable (in statistics)

One measured characteristic that changes from row to row: in a table, one column. Not a Python variable, which is a name holding a value.

e.g. budget is a variable of each project; x = 5 is a Python variable

distribution

How a variable's values are spread out: which are common, which are rare, where they bunch up.

e.g. most budgets are ₱10–70M; a few reach ₱1.45B

mean

The ordinary average: add all the values, divide by how many there are.

e.g. [2, 3, 4, 11] → 20 ÷ 4 = 5

median

The middle value once sorted (the average of the middle two if the count is even). Half below, half above.

e.g. [2, 3, 4, 11] → (3 + 4) ÷ 2 = 3.5

mode

The value that appears most often. The only "typical" that works for words.

e.g. ["M", "L", "M", "S"] → "M"

range

Largest value minus smallest value: the full width the data covers.

e.g. [2, 3, 4, 11] → 11 − 2 = 9

variance

How spread out values are: the average squared distance from the mean (pandas divides by n − 1).

e.g. [2, 4, 6]: mean 4, squared distances 4, 0, 4 → 8 ÷ 2 = 4

standard deviation (SD)

The square root of the variance: how far a typical value sits from the mean, in the column's own units.

e.g. [2, 4, 6] → √4 = 2; in pandas .std()

Words for Today · 2 of 2

Seven words for shape and categories

percentile / quartile

The p-th percentile has p% of the values below it. Quartiles are the 25th, 50th and 75th percentiles: Q1, median, Q3.

e.g. if the 90th percentile of scores is 85, nine in ten scored below 85

IQR

Interquartile range: Q3 − Q1, the width of the middle half. Extreme values barely move it.

e.g. Q1 = 71, Q3 = 88.5 → IQR = 17.5

skew

Lopsidedness. Right-skewed: a long tail of big values (mean above median). Left-skewed: the mirror image. About 0: symmetric.

e.g. incomes are right-skewed: a few people earn far more

histogram

A bar chart of one numeric column: each bar counts how many values fall in one range.

e.g. ax.hist(income, bins=6)

bins

The equal-width ranges a histogram counts into. More bins: narrower bars, more detail, more noise.

e.g. scores 0–100 in 10 bins: 0–10, 10–20, …

box plot

A picture of the quartiles: a box from Q1 to Q3, a line at the median, whiskers, and dots for outliers.

e.g. df.boxplot(column="budget")

frequency table

Each category with how many times it appears (or its share). pandas builds it with value_counts().

e.g. Visayas 6, Mindanao 4, Luzon 2

The Idea

Exploration comes before confirmation

"Exploratory data analysis — John Tukey's phrase — is detective work: looking at the data with as few preconceptions as possible, letting it suggest the questions worth asking, before any hypothesis gets tested."
Describe first. Model later. Never the reverse.
Part A

Centre & Spread

Two numbers that, together, sketch a variable.

You can already load a table and pick one column (weeks 4–6). This part adds the numbers that describe that column: where its middle is (the centre) and how far its values stray from it (the spread).

First Look

describe() — the whole sketch in one call

Count, centre, spread and quartiles for a numeric column. Run it on every column before you believe anything about the data.

scores is a pandas Series (one column of values) holding 12 exam scores. .describe() is a method: a function that belongs to the Series and is called with a dot.

print(scores.describe().round(2)) count 12.00 mean 78.42 std 12.70 min 55.00 25% 71.00 50% 81.00 ← the median 75% 88.50 max 95.00 dtype: float64
  • count 1212 values that are not missing
  • mean 78.42the ordinary average: 941 ÷ 12
  • std 12.70standard deviation: a typical score sits about 12.7 points from the mean
  • 25% / 50% / 75%the quartiles: a quarter of scores are below 71, half below 81, three quarters below 88.5
  • min / maxlowest and highest score: your first check for odd values, for free
Old Friends

You already know these quartiles

The 25% / 50% / 75% rows are the Q1, median and Q3 from week 5's IQR fences — now doing descriptive duty instead of outlier patrol.

Reminder from week 5

quartile
sort the values and cut them into four equal-count groups; Q1 has 25% below it, Q3 has 75% below it
IQR
Q3 − Q1, the width of the middle half: 88.5 − 71 = 17.5 here

50% = the median

Half the values below, half above. A statement about order, so extremes barely move it.

Q1 to Q3 = the middle half

The box in a boxplot. Where the "usual" values live, whatever the tails are doing.

Same numbers, two jobs

Week 5 used them to fence suspects; this week they describe the crowd.

Two "Typicals"

Mean and median agree — until the data leans

On tight, symmetric data they're twins. Add one huge value and the mean chases it while the median stays home.

Seat a billionaire on a bus with 49 teachers. The mean wealth on the bus becomes tens of millions; the median passenger is still a teacher.

tight = pd.Series([10, 12, 11, 13, 12]) print(tight.mean(), tight.median()) 11.6 12.0 ← agree skewed = pd.Series([10, 12, 11, 13, 120]) print(skewed.mean(), skewed.median()) 33.2 12.0 ← diverge
  • pd.Series([10, 12, 11, 13, 12])one column of five values
  • print(a, b)print both numbers on one line: the mean first, then the median
  • 11.6 12.0mean (10 + 12 + 11 + 13 + 12) ÷ 5 = 58 ÷ 5; median: sorted 10, 11, 12, 12, 13, the middle one
  • 33.2 12.0166 ÷ 5; the 120 alone lifts the mean. Sorted, the middle value is still 12

Which to report?

For skewed data — incomes, house prices, wait times — the median is the honest "typical." Better: report both; their gap is itself information.

The Other Half

Same mean, wildly different stories

Standard deviation (SD, std) measures how far values typically sit from the mean. The range is simply largest minus smallest. Two series can share a centre and share nothing else.

calm = pd.Series([50, 51, 49, 50, 50]) print(calm.mean(), round(calm.std(), 1), calm.max() - calm.min()) 50.0 0.7 2 wild = pd.Series([10, 90, 30, 70, 50]) print(wild.mean(), round(wild.std(), 1), wild.max() - wild.min()) 50.0 31.6 80
  • round(calm.std(), 1)the standard deviation, rounded to one decimal
  • calm.max() - calm.min()the range: largest minus smallest
  • 50.0 31.6 80mean, SD, range: the same centre as calm, forty-five times the spread

Why a mean alone misleads

"Average river depth: 1 meter" has drowned people. The mean says nothing about the deep parts — spread does. Next slide: the SD of wild, worked out by hand.

Worked Example

A standard deviation, by hand

Take wild = [10, 90, 30, 70, 50], whose mean is 50. Four steps turn "how far from the mean?" into one number.

ValueDistance from 50Squared
10−401,600
90+401,600
30−20400
70+20400
5000

Squaring stops the minus and plus distances cancelling out; the square root at the end turns the answer back into the original units.

wild = pd.Series([10, 90, 30, 70, 50]) ((wild - wild.mean())**2).sum() 4000.0 wild.var() # 4000 / (5 - 1) 1000.0 wild.std() # square root of 1000 31.622776601683793
  • wild - wild.mean()subtract 50 from every value at once: the distances
  • (…)**2 … .sum()square each distance and add them up: 1,600 + 1,600 + 400 + 400 + 0
  • wild.var()the variance: that total divided by n − 1 = 4 (pandas' choice for a sample)
  • wild.std()√1000 ≈ 31.6: a typical value is about 32 away from 50
On Real Money

describe() on 34,079 budgets

One call, and every summary from Part A at once. Read it top to bottom before you plot anything. df is the flood-control table (34,079 projects); df.budget is a shortcut for df["budget"], its budget column in pesos.

The tell is mean vs 50%

They differ by ₱9.3 million here. On a symmetric distribution (the left half a mirror image of the right) they would be nearly equal. This one leans.

pd.set_option("display.float_format", "{:,.0f}".format) print(df.budget.describe()) count 34,079 mean 46,952,340 std 48,594,319 min 0 25% 14,649,406 50% 37,676,233 75% 67,549,270 max 1,447,499,996 Name: budget, dtype: float64
  • pd.set_option(…)a display setting, as in week 5: show numbers with commas and no decimals
  • mean 46,952,340the average project budget: about ₱47 million
  • 50% 37,676,233the median: half the projects cost less than ₱37.7 million
  • min 0 / max 1,447,499,996some rows record a zero budget; the biggest is ₱1.45 billion
Before You Trust Any Of It

describe() quietly dropped 6,415 rows

Every statistic it printed is about the rows that survived.

The count Line Is The Warning

It is the only place the omission appears

duration is the column you built in week 6: days from start date to completion date. Ask for a summary of it on a 34,079-row frame and pandas reports count 27,664. It did not error, it did not warn — it excluded the missing and carried on.

Which 6,415? The unfinished ones

Exactly the On-Going, For Procurement and Not Yet Started projects from week 5. So "median duration 202 days" is really "median duration of projects that finished".

len(df) 34079 df.duration.describe()["count"] 27664.0 # 6,415 missing # always compare the two round(df.duration.isna().sum() / len(df), 3) 0.188 # 18.8% excluded
  • len(df)counts the rows of the table
  • describe()["count"]picks one line of the summary by its label: how many values were actually used (27664.0 has a decimal because describe() stores every line as a decimal number)
  • isna().sum()isna() marks each missing cell (NaN) True; .sum() counts the Trues, because True counts as 1
  • 0.1886,415 ÷ 34,079: almost one row in five has no duration
Say What The Number Is About

The same statistic, two honest captions

✗ "Flood control projects take a median of 202 days"

Implies all 34,079. It is a statement about a population that was never measured — the unfinished ones have no duration by definition.

Why this one is tempting

It is what .describe() appears to say, and nothing in the output contradicts it except a number you have to notice.

✓ "Of the 27,664 completed projects, the median took 202 days"

States the population, the count and the statistic. A reader can check it and knows what it excludes.

The habit

Put the denominator in the sentence. If you cannot say what it is, you do not yet know what you measured.

Two Typicals, One Column

A 25% gap, and which one you quote is a choice

Spread

A standard deviation bigger than the median

₱48.6M of standard deviation (SD) against a ₱37.7M median is a warning in itself. For bell-shaped data, about 95% of values sit within 2 SDs of the mean (the "mean ± 2 SD" rule). Here that range would reach below zero.

When SD stops helping

It assumes a roughly symmetric distribution. On skewed money, quote percentiles instead — they are always interpretable.

round(df.budget.std()) 48594319 # mean - 2*SD is NEGATIVE 46_952_340 - 2*48_594_319 -50236298 # a budget cannot be negative, so the # symmetric picture is simply wrong here
  • round(df.budget.std())the standard deviation of every budget, rounded to whole pesos
  • 46_952_340the underscores are digit separators, like commas; Python ignores them
  • -50236298about −₱50.2M: mean minus two SDs is a negative budget: impossible, so the bell-shape rule does not apply
Naming The Lean, Numerically

Skew 5.13 is not "slightly right"

Skewness (skew) puts a number on lopsidedness. Positive = a long tail of big values; negative = a long tail of small ones. Roughly: |skew| (the size, ignoring the sign) under 0.5 is symmetric, over 1 is strong.

Kurtosis 76 says: heavy tail

Kurtosis measures how often extreme values turn up. A normal distribution (the symmetric bell curve) scores 0 in pandas. 76 means extreme values are far more common than a bell curve would predict.

round(df.budget.skew(), 2) 5.13 # strongly right-skewed round(df.budget.kurtosis(), 2) 76.01 # very heavy tail # for comparison, a normal sample: bell = np.random.default_rng(0).normal(size=34079) round(pd.Series(bell).skew(), 4) -0.0072 # about zero
  • np.random.default_rng(0).normal(size=34079)NumPy draws 34,079 random numbers from a bell curve (the 0 is a seed, so the draw repeats exactly)
  • round(pd.Series(bell).skew(), 4)wrap them in a Series so pandas can measure the skew, rounded to 4 decimals: about 0, as a symmetric shape should be
Percentiles Beat Mean±SD

Every one of these is a sentence you can say

For skewed data, quote the percentile ladder. The p-th percentile is the value with p% of the data below it. Each rung is directly interpretable without assuming any shape at all.

p99 is the useful "big project" line

₱212M. Above that you are in the top 1% of 34,079 contracts — a much clearer threshold than "two standard deviations".

pd.reset_option("display.float_format") # undo the comma setting p = df.budget.quantile([.05, .25, .5, .75, .9, .99]) print((p / 1e6).round(1)) # in millions of pesos 0.05 1.9 0.25 14.6 0.50 37.7 0.75 67.5 0.90 96.5 0.99 212.1 Name: budget, dtype: float64
  • quantile([.05, .25, …])a list of fractions: .05 asks for the 5th percentile, .5 for the median
  • (p / 1e6).round(1)divide by one million (1e6) and keep one decimal, so ₱96,499,614 reads 96.5
  • 0.90 96.590% of projects cost less than ₱96.5 million
Concentration

Where the ₱1.6 trillion actually sits

Slice of projectsCountShare of total spend
Top 1% by budget3406.5%
Bottom 50% by budget17,03916.5%
The middle16,70077%
Quick Check

Tap to reveal

Two sections took the same exam. Both average 75. Section A's std is 3; Section B's is 18. Who should the teacher worry about?

A · Neither — the averages match
B · Section A — low variety means cheating
C · Section B — same centre, but many students far above and far below 75
D · Both equally, always

C — the spread hides the struggling students.

With std 18, Section B plausibly has students in the 40s that Section A simply doesn't. Identical means, completely different classrooms — which is why a mean without spread is half a sentence.

Break

  Five minutes

Then: what the numbers can't show — shape.

Part B

Shape

Distributions have geography. Draw the map.

You can now sum up a column in a few numbers. Numbers hide the shape; this part draws it, with histograms, box plots and a log scale.

The Workhorse

A histogram shows where values live

Chop the range into bins (equal-width ranges, like 15–20k, 20–25k), count what falls in each, draw bars. Clusters, gaps and tails appear instantly — things no single number can say.

Line people up by height band: how many are 150–155 cm, how many 155–160 cm… The height of each queue is a bar.

fig, ax = plt.subplots() ax.hist(income, bins=6) ax.set_xlabel("Income (k)") ax.set_ylabel("Count") plt.show()
  • fig, ax = plt.subplots()start a blank chart: fig is the whole picture, ax the plotting area you draw on (matplotlib, imported as plt)
  • ax.hist(income, bins=6)income is a Series of 12 incomes (in thousands); count them into 6 equal-width bins, one bar per bin
  • set_xlabel / set_ylabellabel the axes: what is across, what is up
  • plt.show()display the finished chart

Read it in this order

Where's the tallest bar (the crowd)? Is there one hump or two? Which way does the tail stretch?

What That Code Draws

One crowded bar, one lonely income far to the right

Histogram of the 12 incomes in thousands with 6 bars: a tall bar of 10 incomes from 15 to 22.5k, a short bar of 1 income from 22.5 to 30k, an empty stretch from 30 to 52.5k, and a short bar of 1 income from 52.5 to 60k.
Fine Print

Bin count changes the story

The same twelve incomes with bins=6 and bins=3: the outlier (60k) stands alone in both, but with bins=3 all eleven other incomes melt into one bar and the 23k income disappears into the crowd. Neither is "the" picture — both are choices.

Bins are like the size of a fishing net's holes: too big and everything slips into one catch, too small and every fish gets its own pocket.

Too few bins

Everything smooths into one lump; structure (and outliers) vanish.

Too many bins

Every value gets its own lonely bar; noise masquerades as pattern.

The honest practice

Try several bin counts before you conclude anything. If a "finding" only exists at one bin setting, it's probably the bins talking.

What That Code Draws

bins=6 against bins=3, same twelve incomes

Histogram with 6 bins: bars of 10, 1, 0, 0, 0 and 1 incomes, with bin edges every 7.5k from 15 to 60.Histogram with 3 bins: a bar of 11 incomes from 15 to 30k, an empty bin from 30 to 45k, and a bar of 1 income from 45 to 60k.
Draw It

The histogram of the column you just described

Everything above — the lean, the heavy tail, the gap between mean and median — is one picture. Plot it before you believe any of the numbers.

Clip the axis to see the body

With a ₱1.45B maximum, 95% of the data crushes into the first four of 50 bars. Limit the range, and say that you did.

df.budget.hist(bins=50) # → 43% of projects in the first bar (under ₱29M), # then a nearly empty smear out to ₱1.45B # readable version df.budget[df.budget < 2e8].hist(bins=50) # → most projects under ₱50M, spikes near ₱50M # and ₱100M, then a long thin right tail # always note the clip in the caption
  • df.budget.hist(bins=50)pandas' shortcut: draw a histogram of the column with 50 bins (it draws a picture and prints nothing; the # → comments describe what you see)
  • 2e8e-notation: 2 followed by 8 zeros, i.e. 200,000,000 (₱200 million)
  • df.budget[df.budget < 2e8]keep only budgets under ₱200M (a boolean filter: keep the rows where the test is True), then draw those
What That Code Draws

All 50 bars, then the clipped version

Histogram of all 34,079 budgets with 50 bars, x axis from 0 to 1.45 billion pesos (marked 1e9): one bar near zero reaches about 14,500 projects, a few bars follow, and the rest of the axis is almost empty.Histogram of the 33,720 budgets under 200 million pesos with 50 bars, x axis marked 1e8: many bars below 50 million, the tallest spike just under 50 million (about 3,300 projects), a second spike just under 100 million (about 2,400), then a long thin tail to 200 million.
Bins Change The Story

Same 33,000 projects, three different pictures

BinsBin widthTallest bar holdsWhat it looks like
10₱20.0M11,817 projectsOne dominant block — detail gone
30₱6.7M4,891 projectsA clear peak and tail — usable
100₱2.0M3,192 projectsSpiky — noise starts to show
Naming The Lean

Symmetric, right-skewed, left-skewed

When The Shape Fights You

Change the scale, not the data

A log transform does not delete the outliers — it re-spaces the axis.

log₁₀ asks "10 to the power of what?": log₁₀(1,000) = 3, log₁₀(1,000,000) = 6. Each step of 1 means "ten times bigger".

The Log Transform

Skew 5.13 becomes −0.82

Money is often roughly log-normal: its logs look like a bell curve, because differences that matter are multiplicative (×2, ×10) rather than additive (+₱1M). Plotting log₁₀(budget) turns that shape into a near-symmetric one.

Report in pesos, plot in logs

The transform is for seeing. Convert back before you quote a number, or your reader has to exponentiate in their head.

lb = np.log10(df.budget[df.budget > 0]) round(lb.skew(), 2) -0.82 # was 5.13 # the log-mean, back in pesos round(lb.mean(), 2) 7.47 round(10**lb.mean()) 29733984 # geometric mean, ₱29.7M # note: below the median, not above — # the opposite of the arithmetic mean
  • df.budget[df.budget > 0]keep the positive budgets: the log of 0 does not exist
  • np.log10(…)NumPy takes log₁₀ of every value: ₱10M becomes 7, ₱100M becomes 8
  • round(10**lb.mean())** is "to the power of": undo the log. The result is the geometric mean, a typical value on the multiplying scale
Week 5's Rule, Drawn

The boxplot is the IQR rule as a picture

lower fence Q1 median Q3 upper fence outlier whiskers reach the last point inside 1.5 × IQR
Boxplot, On Real Budgets

Week 5’s IQR rule, drawn

The box is p25 to p75 (Q1 to Q3). The whiskers reach 1.5×IQR beyond the box: the top fence is 67.5M + 1.5 × 52.9M = 146.9M. Everything beyond is drawn as a point — the same 798 projects week 5 flagged.

Best for comparing groups

One box per region puts 18 distributions side by side, which a histogram cannot do legibly.

df.boxplot(column="budget", by="region") # box = 14.6M to 67.5M (the IQR) # line = 37.7M (median) # whisker top = 146.9M # dots = 798 projects above it # on skewed data, expect dots. # they are not errors — they are the tail
  • df.boxplot(column="budget", by="region")draw one box of budgets for each region, side by side
  • # box / line / whisker / dotsthe comments list what the national box shows; the numbers come from the describe() and IQR fence above
What That Code Draws

Eighteen boxes, and eighteen labels on top of each other

pandas boxplot of budget by region with the title Boxplot grouped by region, budget; 18 boxes in alphabetical order along the x axis, whose region names overlap into an unreadable strip; the first box, Central Office, is much taller and higher than the rest, and most other boxes are squashed near zero with columns of outlier circles above them; y axis in billions of pesos (1e9).
Choosing

Histogram for detail, boxplot for comparison

They summarize the same distribution at different zoom levels — so they answer different questions.

Reach for the histogram when…

You're meeting a variable for the first time and want its full shape: humps, gaps, tails.

Reach for the boxplot when…

You need a compact summary — especially several side by side: income by region, scores by section.

Next week's bridge

"Boxplots side by side" is already a two-variable question — exactly where Data Exploration II begins.

Quick Check

Tap to reveal

Household income in a city is strongly right-skewed. Without computing anything: how do its mean and median compare?

A · Mean > median — the long right tail pulls the mean up
B · Mean < median
C · Exactly equal, by definition
D · Impossible to say anything

A — the tail drags the mean.

A few very high incomes lift the mean well above what a typical household earns; the median stays with the crowd. It's why "average income" headlines flatter and "median income" informs.

Part C

Categories & The Habit

Variables with no mean at all — and the checklist that covers everything.

You can describe a column of numbers. Text columns such as region or contractor have no mean; this part counts them, then gathers every check into one routine.

Describing Categories

You can't average a region — you count it

A categorical column holds labels, not amounts (region, status). Its summary is a frequency table: each value and how often it appears. value_counts() builds it, and with normalize=True gives each value's share.

r = pd.Series(["Visayas"] * 6 + ["Mindanao"] * 4 + ["Luzon"] * 2, name="region") r.value_counts() region Visayas 6 Mindanao 4 Luzon 2 Name: count, dtype: int64 r.value_counts(normalize=True).round(2) region Visayas 0.50 Mindanao 0.33 Luzon 0.17 Name: proportion, dtype: float64
  • r = pd.Series([…] * 6 + …)a toy column of 12 region labels: * 6 repeats a label six times, + joins the lists
  • r.value_counts()count the rows for each region, biggest first: 12 rows in total
  • normalize=Truedivide each count by the total: 6 ÷ 12 = 0.50. The shares add up to 1 (proportion = share)

Counts vs shares

Counts answer "how many"; shares answer "how dominant" — and shares survive comparison across datasets of different sizes.

Categories

On real data: 18 regions, 4,841 contractors

The whole Part A toolkit is for numbers. For text columns, the summary is a frequency table, and the "typical" value is the mode: the value that appears most often.

Check unique first

18 regions is groupable. 4,841 contractors is a long tail. 33,325 different descriptions is free text — a different tool entirely.

df.region.value_counts(normalize=True).round(3).head(4) region Region III 0.159 National Capital Region 0.115 Region I 0.103 Region IV-A 0.101 Name: proportion, dtype: float64 df.region.nunique(), df.contractor.nunique() (18, 4841) # mode = the most common value df.region.mode()[0] 'Region III'
  • .head(4)show only the first four lines: Region III holds 15.9% of projects
  • nunique()how many different values the column has
  • mode()[0]mode() returns a list-like Series, since two values can tie; [0] takes the first
Your Turn · 8 min

Describe one column properly

1 · Pick a number column

Use duration (you built it in week 6) or progress (percent complete, 0–100). Run .describe().

2 · Measure the lean

Compute .skew(). Predict the histogram shape from that number alone, then plot it and check.

3 · Try three bin counts

10, 30, 100. Which one would you put in a report, and what is your reason?

4 · Write three sentences

One about the centre, one about the spread, one about the shape — each with a number in it.

Quick Check

Tap to reveal

Flood budgets have mean ₱46.95M, median ₱37.68M and skew 5.13. Which sentence is the most honest summary?

A · "The average flood control project costs ₱47 million"
B · "Half of projects cost under ₱38 million; the distribution is strongly right-skewed"
C · "Projects cost ₱47M ± ₱49M"
D · "Most projects cost close to ₱47 million"
B. A is arithmetically true but suggests a typical project that does not exist. C implies symmetry — mean minus 2 SD is negative here, which is impossible for a budget. D is simply false: the mode is well below the mean. With skew 5.13, quote the median and say the distribution leans.
Quick Check

Tap to reveal

You take log₁₀ of the budget column and the skew goes from 5.13 to −0.82. What have you achieved?

A · Removed the outliers from the dataset
B · Re-spaced the axis so the shape is visible — every row is still there
C · Corrected an error in the data
D · Made the mean equal the median
B. A log transform changes the scale, not the data. All 34,079 projects remain, in the same order. It works here because money tends to be log-normal — what matters is multiplicative. Plot in logs, but convert back to pesos before quoting any figure.
The Routine

The univariate EDA checklist

The ECDF

"What fraction is below X" — no binning required

A histogram needs you to choose bins. The ECDF (empirical cumulative distribution function) does not: for any threshold you pick, it gives the share of the data below it.

Read it as a sentence

Each row below is a claim you could put in a report, with no modelling assumption behind it.

round((df.budget < 1e7).mean(), 3) 0.214 # 21.4% under PHP 10M round((df.budget < 5e7).mean(), 3) 0.709 # 70.9% under PHP 50M round((df.budget < 1e8).mean(), 3) 0.936 # 93.6% under PHP 100M # 6.4% of projects are over PHP 100M # and they hold 22.7% of all the money
  • df.budget < 1e7compare every budget with 1e7 = 10,000,000: one True or False per row
  • ( … ).mean()True counts as 1 and False as 0, so the mean is the share of Trues: 0.214 = 21.4% of projects
Comparing Spread Across Units

CV puts pesos and days on the same ruler

A standard deviation of 48 million pesos and one of 159 days cannot be compared. The coefficient of variation (CV) divides each SD by its own mean, so the units cancel and both become plain numbers.

Rule of thumb

CV above 1 means the spread exceeds the average — a sign the distribution is skewed or has heavy tails. Budget sits at 1.03.

cv = lambda x: x.std() / x.mean() round(cv(df.budget), 2) 1.03 # sd exceeds the mean round(cv(df.duration), 2) 0.68 # much tighter # projects vary more in COST # than in how long they take
  • cv = lambda x: x.std() / x.mean()lambda makes a one-line function (you met it in week 5): give it a column x, it hands back SD ÷ mean
  • cv(df.budget)call it on the budget column: 48.6M ÷ 47.0M = 1.03
Same Statistic, Different Regions

Where is spending consistent, and where is it lumpy?

RegionProjectsMedian budgetCV
Region IV-B1,467₱37.6M1.18 — lumpiest
Region VIII1,935₱27.5M1.07
Central Office180₱144.5M1.05
Region III5,412₱46.3M0.72 — most consistent
The Routine, Applied

Six lines that describe any numeric column

Run these in order on every new column. It takes a minute and it is the difference between describing data and guessing at it. Here col stands for any one column (for example df.budget) and threshold for a cut-off number you choose.

Then, and only then, plot

The numbers tell you what shape to expect. If the picture disagrees, one of you is wrong and it is worth finding out which.

col.isna().sum() # what is missing col.describe() # centre + spread col.skew() # which way it leans col.quantile([.05,.95]) # the realistic range (col < threshold).mean() # ECDF at a point col.hist(bins=30) # look at it # six lines. every column. every time.
Why Pictures Are Not Optional

Identical statistics, different data

"Anscombe's quartet: four datasets with the same means, the same variances, even the same correlation — and four completely different shapes when plotted. The summary numbers cannot tell them apart. A single glance can."
F. J. Anscombe, 1973 — the most famous argument for drawing your data
Glossary

Every new word from today, in one line each

Centre and spread

EDA
exploratory data analysis: summaries and charts before testing ideas
variable (in statistics)
one measured characteristic, i.e. one column (not a Python name)
distribution
how a variable's values are spread out
mean
add the values, divide by how many
median
the middle value once sorted
mode
the most frequent value; works for categories
range
largest minus smallest
variance
average squared distance from the mean
standard deviation (SD)
square root of the variance; typical distance from the mean
percentile / quartile
value with p% below it; quartiles are p25, p50, p75
IQR
Q3 − Q1: width of the middle half
coefficient of variation (CV)
SD ÷ mean; compares spread across units
denominator
the "out of how many" behind a number

Shape and categories

skew
lopsidedness; right-skewed has a long tail of big values
kurtosis
how often extreme values appear (normal curve = 0)
normal distribution
the symmetric bell curve
histogram
bars counting values per range
bins
the equal-width ranges a histogram counts into
box plot
box Q1–Q3, median line, whiskers, outlier dots
outlier
a value far from the rest (beyond the 1.5 × IQR fence)
log transform
replace x by log₁₀(x) to re-space a skewed axis
ECDF
share of values below any threshold you pick
categorical
a column of labels, not amounts
frequency table
each category with its count or share
survivorship bias
judging only the cases that got through a filter
This Week's Lab

Describe, compare, draw

Exam scores through describe(), a tight and a skewed series to compare centres, calm-vs-wild spreads, and an income histogram whose story changes with the bins. ~45 minutes.

You'll practise

describe(), mean()/median(), std() and range, and ax.hist() with labeled axes.

Stretch, if you're quick

value_counts(normalize=True) for a categorical, and a boxplot you connect back to week 5's IQR fences.

Recap

Four things to carry out

Readings

Before next week

Next Week

Data Exploration II

Multivariate EDA and visualization — correlations, group comparisons, and the plots that reveal how variables move together.

DS 227 · Knowledge Discovery in Data