Turning a summarised table into a chart that makes the pattern obvious — with matplotlib for control and seaborn for speed.
Programming for Data Science · University of the Philippines Cebu
The fig, ax foundation every Python chart is built on.
Line, bar, scatter, histogram — and when each one fits.
Statistical plots straight from a DataFrame, and honest design.
A chart is how your analysis reaches other people. A good one answers a question before they finish reading the title.
matplotlib is a blank canvas and a set of pens: you draw every part yourself. seaborn is a chart template: hand it a table and say which columns go where.
Make a labelled, honest chart from a DataFrame.
The whole picture: the canvas you size and save as an image file.
e.g. fig.savefig("chart.png")
One plotting area inside a figure, with its own x and y axis. You draw on it.
e.g. fig, ax = plt.subplots() gives one figure and one axes
To draw data onto an axes; ax.plot draws points joined by a line.
e.g. ax.plot([2021, 2022], [12, 15])
The symbol drawn at each data point: a dot, a square, a triangle.
e.g. marker="o" puts a dot on every point
Text that names an axis or the chart, ideally with the unit.
e.g. ax.set_ylabel("Sales (₱)")
The small key that says which colour or marker means which group.
e.g. ax.legend() after label="Cebu"
One of several small charts arranged in a grid inside one figure.
e.g. plt.subplots(1, 2) makes two side by side
Bars that count how many values fall in each range (bin) of one column.
e.g. ax.hist(scores, bins=5)
The basic Python charting library: you draw and label every part yourself.
e.g. import matplotlib.pyplot as plt
A library built on matplotlib that draws statistical charts from a DataFrame in one call.
e.g. sns.histplot(data=df, x="pop") (Colab, not the browser lab)
Your groupby gave numbers per region. But "which is biggest, and by
how much" leaps out of a bar chart far faster than out of a column of digits.
Exact, but you must read every row to compare.
Approximate, but the pattern is instant.
The foundation under every Python chart.
You can already turn a table into summary numbers. This part draws them: matplotlib is the Python library (a package of ready-made code you import) that draws charts, one instruction at a time.
plt.subplots() hands back a figure (the canvas)
and an axes (the plot itself). You call methods on ax to draw and
label.
ax.plot draws points joined by a lineA method (week 4) is a function that belongs to an object, called with a dot: ax.plot(…).
import matplotlib.pyplot as pltload matplotlib's drawing tools under the short name pltfig, ax = plt.subplots()make a new figure with one axes; the function gives back two things, so two names on the left catch themax.plot(years, sales)years and sales are two equal-length lists: x values and y values. Draw a line through the pointsax.set_title("Sales over time")write a title above the plot
sales dip in 2023, then climb. The x axis shows 2021.5 and 2022.5 because matplotlib treats years as ordinary numbers, and neither axis says what it measures yet: labels come a few slides on.
Run with the week-8 lab’s lists: years = [2021, 2022, 2023, 2024, 2025], sales = [12, 15, 14, 20, 24]. Every chart in this deck is the real output of the code on the slide before it.
plt.plot() draws onto whatever figure is "current", which is fine until there are two. Ask for the objects explicitly and that ambiguity never arises.
Everything you save or size happens on fig. Everything you draw or label happens on ax.
figsize=(8, 5)the figure's width and height, in inchesax.scatter(df.budget, df.duration, alpha=0.05)one dot per project: budget across, duration up. duration is a column made first: days from start to completion, (df.completion_date - df.start_date).dt.days after converting both to dates (week 6). alpha is transparency, from 0 (invisible) to 1 (solid)ax.set_xlabel / ax.set_ylabelthe text under the x axis and beside the y axis
even at alpha=0.05 the left edge is solid blue. Most budgets are small and a few are huge, so nearly every dot is squeezed into the first fifth of the axis. The 1e9 at the bottom right means the ticks are in billions of pesos (0.2 = ₱200 million).
Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
Two numbers give you a grid. axes is then a NumPy array, and you index it like one. Each small chart in the grid is called a subplot.
Panels with different y-limits look comparable and are not. Sharing the axis is what makes a multi-panel figure honest.
plt.subplots(2, 2, …)2 rows by 2 columns of subplots; sharey=True gives them all the same y scaleaxes[0, 0]the subplot in row 0, column 0 (top left); counting starts at 0, as with listshist(…, bins=30)a histogram with 30 bars (see Part B)for ax, col in zip(axes.flat, cols)walk through the subplots and a list of column names in pairs (zip pairs them up), drawing one column per subplot
budget piles into the first one or two bars: most projects are small, a few are enormous. Because of sharey=True both histograms use the same 0–22,000 scale, so their bar heights can be compared directly.
Only the two hist lines and tight_layout(): the loop then fills the empty panels with whichever columns cols lists. Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
Overplotting is when dots pile on top of each other until you can no longer see how many there are. Every real scatter in this course has too many points. Three fixes, and the first one is usually enough.
Transparency shows density for free. When even that saturates, bin the plane instead of drawing points.
alpha=0.05, s=4each dot 95% see-through and small (s = size), so dark areas mean many dotsdf.sample(2000, random_state=42)pick 2,000 random rows; the fixed random_state gives the same 2,000 every runhexbin(…, cmap="Blues", xscale="log")split the plane into small hexagons and colour each by how many points fall in it (cmap = colour scale); xscale="log" bins on a log axis, which has no place for a 0 budget, so drop those rows firstset_xscale("log")a log scale: each step along the axis multiplies (1M, 10M, 100M), so tiny and huge budgets both fit. Calling it after hexbin leaves the chart empty
alpha=0.05, s=4, then set_xscale("log")
hexbin(…, xscale="log"), budgets above 0once the budget axis is logged, both show bigger budgets tending to take longer, and budgets bunching at round numbers (the vertical stripes). The darkest hexagon is the most common kind of project.
Here x is df.budget and y is df.duration. Run as first written, hexbin followed by set_xscale("log") drew an empty chart, so the code now passes xscale="log" to hexbin; it refuses the 939 projects with a ₱0 budget, so those are dropped first.
Title, axis labels with units, and the data. Miss any of the first three and a reader has to guess what they're looking at.
A label is the text that names an axis or the chart; the small numbers along each axis are ticks.
Units belong in the axis label: "Sales (₱)", not just "Sales".
Three lines turn a mystery plot into a clear one. Do it every time — future you, reading this in a report, will not remember what the axes meant.
A chart without labels is a map without a legend or place names: the shapes are there, but nobody can say what they mean.
ax.set_xlabel("Year")names the horizontal axisax.set_ylabel("Sales (₱)")names the vertical axis, with its unit in bracketsax.set_title(…)the headline above the chart: what, where, whenMatch the chart to the question.
You now know how to make a figure and label it. This part is about choosing which picture to draw, because each chart type answers a different kind of question.
| Chart | Shows | Reach for it when |
|---|---|---|
| Line | a trend over time | the x-axis is ordered — dates, steps |
| Bar | compare categories | a value per named group |
| Scatter | a relationship | two numbers, one point per row |
| Histogram | a distribution | the spread of one number |
The wrong chart hides the answer. A line chart over unordered categories, for instance, implies a trend that isn't there.
A distribution is how the values of one column are spread out: where most of them sit, and how far the extremes reach. A relationship is whether two numbers tend to rise or fall together.
"Am I showing a trend, a comparison, a relationship, or a spread?"
A line connects ordered points — it says "these flow into each other." Bars sit apart — they say "compare these separate things."
A line chart is a temperature graph through the day: the hours follow each other. A bar chart is a height chart for a class photo: each bar is a separate person.
ax.plot(years, sales)x = years (ordered), y = sales; points joined by a lineax.bar(regions, pops)one bar per region name, as tall as its populationA line chart says "these points are connected in order". Only use it when they are — time, or another genuine sequence.
A marker is the symbol drawn at each data point: "o" a dot, "s" a square, "^" a triangle.
Connecting 18 regions implies Region I flows into Region II. Bars, for anything unordered.
df.groupby("year").budget.sum() / 1e9week 7: total budget per year, divided by 1e9 (one billion) so the numbers read as billions of pesosax.plot(y.index, y.values, marker="o")years across (the Series' index), totals up (its values), a dot on each year... hides 2017–2023): it climbs to ₱369.3B in 2024, then drops, to 0.0 in 2026
the line climbs to a 2024 peak, then plunges to ₱196B in 2025 and to zero in 2026. The chart draws that plunge as confidently as every other year; the next two slides ask whether it is real.
Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
A 47% collapse in flood control spending — or an export taken partway through the year.
2026 has one project in it. 2025 has 3,475 against 2024’s 5,553. The dataset was exported mid-period, so the last bars are partial — and nothing in the chart says so.
Drop the incomplete periods, or draw them differently (dashed, greyed) and label them. Never plot a partial period as if it were whole.
df.groupby("year").size().tail(3)how many projects each year has; tail(3) shows the last three yearsy[y.index <= 2024]keep only the years up to 2024 (the years are whole numbers, so no quotes: "2024" in quotes would raise a TypeError)y[:-2], y[-3:]y[:-2] is every year except the last two, drawn solid; y[-3:] is 2024 to the end, drawn as a grey dashed line, so the drop into the partial years looks different (dashing only y[-2:] would leave the 2024→2025 drop solid)ax.annotate("partial year", ...)write a note on the chart (the ... stands for its position, shown later)
y[y.index <= 2024]
on the left the partial years are gone, so the chart ends at the 2024 peak. On the right they stay but look different: the grey dashes say “not comparable” before anyone reads a label.
The note was placed with xy=(2025, y[2025]), xytext=(2025.3, 250). Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
Three decisions turn an unreadable bar chart into a readable one, and none of them is about colour: sort the bars by size, lay them sideways, and start the value axis at zero.
"National Capital Region" rotated 45 degrees is unreadable and eats a third of the figure. barh costs nothing.
.sort_values()order the regions from smallest to largest total; barh draws the first one at the bottom, so the biggest bar ends up on top and the chart reads as a rankingax.barh(g.index, g.values)barh = horizontal bars: region names down the side, where long names fit
every region name is readable, the longest bar sits on top, and Region III’s bar is about 1.7 times NCR’s, which the eye sees at once without reading a number.
Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
twinx() ("twin x") puts two scales on one plot: a second y axis on the right that shares the same x axis. It is the easiest way to imply a relationship that is not there, because you choose both scales.
Two stacked panels with sharex=True. The comparison survives; the implied correlation (the suggestion that the two lines move together) does not.
ax2 = ax.twinx()a second axes drawn on top of the first, with its own y scale on the rightfig, (a1, a2) = plt.subplots(2, 1, sharex=True)two subplots stacked (2 rows, 1 column) that share the same x axis; the brackets catch the two axes by name
twinx(): two scales on one plot
sharex=Trueon the left the lines seem to rise and fall as one, but only because matplotlib stretched each scale to fill the box; a different right-hand scale would pull them apart. On the right each measure keeps its own axis.
counts = projects per year, df.groupby("year").size(); totals = the yearly budget y; the twin axes sit on the spend-per-year chart. Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
A scatter plots one point per row to reveal whether two numbers move together. A histogram bins one column to show where values cluster.
A histogram cuts the range of values into equal slices called bins (e.g. 0–10, 10–20, …) and draws one bar per bin, as tall as the number of values that fall in it.
A histogram is lining students up by height band: how many are 150–155 cm, how many 155–160 cm, and so on.
ax.scatter(area, pop)one dot per city: area across, population upax.hist(pop, bins=10)only one column is needed; matplotlib counts how many cities fall in each of 10 population bandsYou want to see the spread of exam scores across one class. Which chart?
C — histogram.
A histogram bins one column to show its distribution — where scores cluster and how wide they spread. A line implies order; a scatter needs two numbers.
Spread of one number → histogram. It's the go-to for "what does this column look like?"
One column's shape → histogram.
Back for the fast, good-looking way to plot.
Statistical charts, straight from a DataFrame.
You can build any chart in matplotlib line by line. seaborn is a second library, built on top of matplotlib, that draws common statistical charts from a DataFrame in one call. (The in-browser lab has matplotlib only; try seaborn in Colab.)
Pass the whole frame and name the columns for each axis. seaborn reads the data directly, picks sensible defaults, and labels the axes for you.
With matplotlib you hand over two lists of numbers. With seaborn you hand over the whole table and say "region goes across, pop goes up" by column name.
import seaborn as snsload seaborn under its usual short name snssns.barplot(data=df, x="region", y="pop")a bar per region, height = population; column names go in quotessns.scatterplot(…)a dot per city: area across, population upBecause half of seaborn returns a figure, not an axes.
This one distinction explains nearly every seaborn error message a beginner meets.
ax=The *plot functions that end in plot and take ax= are axes-level. relplot, displot, catplot are figure-level.
ax=ax"draw on this axes": the one you made with plt.subplots()col="region", col_wrap=4one panel per region, four panels per rowg.fig / g.axesg is the grid seaborn built; these give you its figure and its axesrelplot(…, ax=ax)if you pass ax= anyway, seaborn prints a warning that relplot does not accept it and ignores it: the chart lands in a new figure, not on your axesIf you want one chart placed inside a layout you control, use axes-level. If you want the same chart repeated per group, use figure-level and let it build the grid.
There is no col= on scatterplot. Reaching for it and failing is the usual reason people end up writing manual subplot loops.
sns.relplot(…, col="region", col_wrap=6)one budget-vs-duration scatter per region, six panels per row (18 panels)height=2.2, alpha=0.3each panel 2.2 inches tall; dots partly see-throughg.set(xscale="log")apply a log x scale to every panel at onceg.savefig("by_region.png", dpi=150)save the whole grid as an image file
every panel shares the same axes, so regions compare at a glance: Region III and NCR are dense, Central Office has only a handful of dots. Two long region names collide in the bottom row; that is seaborn’s default output.
d is the 2,000-row sample from the overplotting slide, df.sample(2000, random_state=42). Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
| Function | Draws | Extra it gives you |
|---|---|---|
sns.histplot | a histogram | optional density curve (a smooth line tracing the bars’ shape) |
sns.boxplot | spread per group | median + outliers at a glance |
sns.scatterplot | a relationship | hue= colours by a category |
sns.barplot | a value per group | a confidence bar (a thin line showing how uncertain each average is) automatically |
Same charts as matplotlib, fewer lines, better defaults. Reach for seaborn first; drop to matplotlib for fine control.
A box plot draws each group's middle half of
values as a box, with a line at the median and dots for unusual values (outliers).
hue= means "colour by this column".
hue="region" colours points by group with zero extra work.
A legend is the small key that says which colour or marker means what. A series is one set of data drawn as one line or one set of dots. A legend on a single-series chart is noise. On a multi-series one it is essential — and its placement matters more than people expect.
Two or three series read better annotated at their right-hand end than via a legend the eye has to travel to.
label="Region III"give each line a name; nothing shows yetax.legend(loc="upper left", frameon=False)draw the key from those names, in the top-left corner, with no box around itbbox_to_anchor=(1.02, 1)place the key just outside the right edge of the axes (1.0 = the edge)
loc="upper left", frameon=False
bbox_inches="tight"inside, the key fits the empty top-left corner. Outside, a default save cuts the names off at the image edge; saving with bbox_inches="tight" brings the whole key back.
x = the years; a, b = Region III’s and NCR’s yearly budget in PHP billions, from df.groupby(["year", "region"]).budget.sum(). Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
Consistent figures across a report come from one call, not from remembering to pass the same arguments twenty times.
whitegrid for anything with a y-scale to read against; ticks for scatter plots where a grid adds nothing.
style="whitegrid"white background with light grid linespalette="colorblind"the set of colours used for groups, chosen to stay distinguishable for colour-blind readerscontext="talk"scale all text and lines for a projector| Symptom | Cause | Fix |
|---|---|---|
| Nothing appears in the notebook | No plt.show(), or a non-interactive backend | Add %matplotlib inline |
| Labels cropped in the saved file | Default save clips outside the axes box | bbox_inches="tight" |
relplot ignores ax= (with a warning) | It is figure-level and owns its figure | Use scatterplot(ax=ax) |
| Panels overlap each other | Default spacing does not account for labels | fig.tight_layout() |
All four are layout problems, not data problems — which is why they are worth memorising. The chart was right; the frame around it was not.
%matplotlib inline is a notebook command (not Python) that shows charts under the cell. A backend is the part of matplotlib that turns a figure into pixels on a screen or in a file.
A cropped y-axis turns a 2% gap into a cliff. Bars must begin at 0.
One clear message per chart beats ten series fighting for attention.
Colour should encode meaning, not decorate. Keep it readable for everyone.
The echo from DS 227 (the companion course): presenting data is a choice. The axis you pick shapes the story a reader takes away.
If a design choice inflates the effect, it's a lie by chart.
A bar’s meaning is its length. Cut the axis and you have changed every length while keeping the numbers technically correct.
A line chart encodes position, not length, so a non-zero baseline is fine — and often necessary to see the change at all.
ax.set_ylim(0, 1)the y axis runs from 0 to 1 (0% to 100%)ax.set_ylim(0.6, 0.95)the y axis now starts at 60%, so every bar loses its bottom 60% and the gaps look far bigger
set_ylim(0, 1)
set_ylim(0.6, 0.95)on the left the bars differ modestly, which is true. On the right the same numbers make Region XII look almost six times Region IX. The names crashing along the bottom are the long-label problem that barh solves.
rates = each region’s share of projects with status Completed, sorted. Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
On a log axis (short for logarithmic), equal distances are equal ratios, not equal amounts. That is a completely different chart, and it must be labelled as one.
Not just the caption — the label travels with the image when someone screenshots it into a slide.

the same dots as the first scatter now fill the width. The ticks go 106, 107, 108, each ten times the last, and the axis label says so.
The first scatter (figsize=(8, 5), alpha=0.05) with these two lines added. Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
About 8% of men have some form of colour-vision deficiency. Red-green is the most common failure and the most common default.
A colormap (cmap) is a colour scale that turns numbers into colours, like a heat map from pale to dark.
Sequential for magnitude, diverging for a meaningful midpoint, qualitative for unordered categories. Using the wrong family misstates the data.
center=0)A reader should not have to find your conclusion in the dots. Mark it.
Two arrows mean the chart has two messages, which usually means it should be two charts.
xy=(x0, y0)the data point the arrow points atxytext=(x0*2, y0+50)where the note's text sits, away from the pointarrowprops=dict(arrowstyle="->")draw a simple arrow from the text to the point"\n" in the texta line break: the note is written on two lines
the title states the finding and one arrow points at the one group being discussed, so a reader gets the message without hunting through the dots. The note runs past the right edge because xytext doubles x0; a tight save keeps it in the file.
x0, y0 = Central Office’s median budget (₱144.5M) and median duration (448 days), on the log-scale scatter from the previous slides. Data: DPWH flood-control projects (BetterGov.ph mirror, CC0), 34,079 rows; 27,664 have a duration.
tight_layout() stops labels clipping; savefig with a
higher dpi exports a crisp image for a report or slide.
dpi means "dots per inch": more dots give a sharper, larger image file.
fig.tight_layout()nudge labels and subplots so nothing overlapsfig.savefig("pops.png", dpi=150)write the figure to an image file called pops.png next to your notebookplt.show()display the chart on screenThe default save clips anything outside the axes box — which is exactly where rotated tick labels and legends live.
150 for a slide, 300 for print. The default 100 looks soft on any modern screen.
bbox_inches="tight"crop the saved image to everything drawn, including labels outside the axes"fig.pdf"a vector file: stored as shapes, not pixels, so it stays sharp when zoomedplt.close(fig)throw the figure away once saved, to free memory| Check | Why it matters |
|---|---|
| Both axes labelled, with units | "Budget" is not a unit. "Budget (PHP)" is. |
| Title states the finding | "Larger projects take longer" beats "Budget vs Duration". |
| Bars start at zero | Length is the encoding; truncating it lies about ratios. |
| Log scales declared in the label | Otherwise the reader assumes linear distances. |
| n stated somewhere | "27,664 projects with a completion date" — the denominator is part of the claim. |
| Saved with bbox_inches="tight" | Or the labels you just wrote get cropped off. |
Six checks, about a minute. It is the difference between a chart that supports your argument and one that someone has to interrogate you about.
Scatter budget against duration with subplots(). Fix the overplotting with alpha, and log the x-axis.
Both axes with units, a title that states a finding rather than naming the variables.
Rebuild as relplot with col="region", col_wrap=6. Confirm the axis limits are shared.
PNG at dpi=150 with bbox_inches="tight". Open the file and check nothing is cropped.
Step 4 catches more problems than any other step. Charts that look fine inline are cropped surprisingly often once saved.
You call sns.relplot(data=d, x=..., y=..., ax=ax) and you get a warning, and your axes stay empty. Why?
scatterplot, boxplot, histplot) draw onto an axes you supply. Figure-level ones (relplot, displot, catplot) build and own their whole figure — which is what lets them facet with col=. Use scatterplot(..., ax=ax), or drop ax= and use g.fig afterwards.Completion rates across regions run from 65.6% to 92.0%. You set ax.set_ylim(0.6, 0.95) on a bar chart. What have you done?
Summarise, plot as bars, and label. Group, chart, label, save — the whole pipeline from a table to a report-ready figure.
df.groupby("region")["pop"].sum()week 7: total population per region, as a Seriesby_reg.plot.bar()pandas' shortcut to matplotlib: draw a bar chart of the Series and give back its axesax.set_ylabel / ax.set_titlethe same labelling as Part A, on that axesA bar chart's y-axis starts at 90 instead of 0. What's the risk?
B — it exaggerates small differences.
Bar length encodes value, so a cropped baseline makes a tiny gap look huge. Bars must start at zero; a line chart may crop, since it encodes position, not length.
Bars encode length — always from zero. This is the most common honest-chart slip.
Bar baseline = 0, always.
You'll build a line, bar, scatter, and histogram from a real dataset with matplotlib and seaborn, label each properly, and export one to a file. ~45 minutes.
plt.subplots, the four chart types, axis labels, and
sns one-liners.
The in-browser lab runs matplotlib only (seaborn is not installed there), so the seaborn one-liners are for Colab.
Use plt.subplots(1, 2) to place two charts side by side.
fig, ax = plt.subplots(), then draw and label on ax.
Line (trend), bar (compare), scatter (relate), histogram (spread).
One-line stats plots from a DataFrame — and keep them honest.
One sentence: match the chart to the question, always label the axes, and never let the design overstate the data.
Turn any summary into a chart a reader trusts at a glance.
ax.plot joins points with a line"o", "s")twinx); easy to mislead withax=) / builds its own figureFinish the Week 8 lab and submit it. Make sure every chart you make has a title and labelled axes.
seaborn "An introduction to seaborn" — seaborn.pydata.org/tutorial/introduction.
Everything here is linked on the course page beside this deck.
Next week: where data comes from — files, SQL databases, and web APIs.
Beyond the CSV: reading JSON and Excel, pulling rows from a SQL database, and fetching live data from a web API.
DS 208 · Programming for Data Science