DS 227 · Week 10

Data Storytelling

Narrative, dashboards and communication — carrying your finding to an audience with thirty seconds and no context.

Knowledge Discovery in Data · University of the Philippines Cebu

Session Map

From finding to felt

Words for Today · 1 of 2

Eight words for telling the story

audience

The particular person or group who will read your chart. Decide who, before you write anything.

e.g. a fellow analyst, a newspaper editor, the general public

narrative arc

The order a story moves in. Here: three beats, context → contrast → consequence.

e.g. "spend grew" → "48 firms hold a quarter" → "worth asking why"

headline (finding-title)

A chart title written as a full sentence that states what the chart shows, with its number.

e.g. "Luzon sales grew 35% over the half-year", not "Sales"

annotation

A short note written on the chart itself, next to the mark it explains, often with an arrow.

e.g. an arrow to the dip labelled "enrolment period"

chart choice

Picking the kind of chart from the question: line for change over time, sorted bars to compare groups…

e.g. "how did it change?" → a line chart

pre-attentive attribute

A visual property the eye notices in a split second, before reading: colour, size, position, length.

e.g. one orange bar among grey ones jumps out at once

decluttering

Deleting everything on a chart that does not help the message: heavy grid lines, borders, 3-D, extra decimals.

e.g. ₱267.7B instead of ₱267,688,008,746.34

colour meaning

Using colour on purpose: one colour per thing, the same in every chart, and grey for "background".

e.g. Visayas is orange in every panel of the dashboard

Words for Today · 2 of 2

Eight words for building the dashboard

dashboard

One screen of a few charts and headline numbers that people check again and again to answer one question.

e.g. a page showing spend per year, by region, and the top contractors

KPI (key performance indicator)

One headline number a dashboard tracks, shown big so it is read first.

e.g. "₱1.60T total, 11 years" in the top-left tile

small multiples

A grid of small charts side by side, each making one point, so the eye can compare them.

e.g. a trend panel next to a totals panel

filter

A control that narrows the data a dashboard shows, such as a drop-down list or a date slider.

e.g. pick "Region VII" and every chart redraws for Region VII only

interactivity

A chart that responds to the reader: hover for exact values, click to filter, zoom in.

e.g. hovering a bar shows "Luzon: 134"

pivot table

A reshaped table: one row per value of one column, one column per value of another.

e.g. rows = months, columns = regions, cells = sales

figure & axes

In matplotlib, the figure is the whole picture; each axes is one chart panel inside it.

e.g. plt.subplots(1, 2) gives one figure holding two axes

legend

The small key on a chart that says which colour or line style stands for which thing.

e.g. an orange line = Visayas, a blue line = Luzon

Part A

Narrative

The reader starts from zero. Design for that.

You already know how to find and check one finding (weeks 7–9). This part adds the next step: choosing who you tell it to and in what order.

The Gap

You know too much to see your own chart

You've lived in this data for weeks — every column name, every caveat. Your reader gets one screen and thirty seconds. The curse of knowledge is forgetting that gap exists: once you know something, it is hard to imagine not knowing it.

It is like giving directions to your house as "turn left where the old bakery used to be". Obvious to you; useless to a visitor who never saw the bakery.

What you see

"Obviously the June dip is the enrollment period — everyone knows that."

What they see

An unexplained dip, a legend with cryptic labels, and no reason to care.

The fix

Put the context on the chart: a finding-title, a label at the dip, units on the axis. Assume nothing survives outside the image.

Same Finding, Three Audiences

The number does not change; everything else does

"48 of 4,841 contractors hold 27.6% of awarded value" is one fact. Your audience, the particular reader you are telling, decides its length, its framing, and what you lead with.

Words in the analyst version

n=4,841
n is how many things were counted: here, 4,841 contractors
top 1%
the 48 contractors with the largest summed budgets (1% of 4,841)
Gini-adjacent
close to the Gini index, a 0–1 score of how unequal shares are (0 = all equal)
conduct claim
an accusation that someone did something wrong

Write the audience down first

Most unclear communication is a piece written for nobody in particular, which means it is tuned for the writer.

# to a fellow analyst "Top 1% of contractors by summed budget = 27.6% of awarded value, n=4,841. Gini-adjacent; no conduct claim." # to an editor "A small group of firms holds a quarter of flood control money. Structural, not an allegation." # to the public "48 companies. A quarter of the budget."
What Each Audience Needs

Three different deliverables from one analysis

The So-What Test

If you cannot finish the sentence, it is not a story yet

Take any finding and say it out loud, then answer "so what?" twice. (The median on the right is week 7's middle value: half the projects took longer, half shorter.) If you run out before the second answer, you have a statistic, not a story.

A worked example

It usually takes two hops to reach something a reader cares about — and the second hop is where you find out whether your evidence reaches.

"Region VI projects take a median 260 days; Region I takes 167." # so what? "Nearly 100 days of difference in comparable flood control work." # so what? "Communities in Region VI wait three months longer for the same protection — worth asking why." # THAT is the sentence you lead with. # and note it is a question, because # the data cannot answer the why
Structure

Every data story has three beats

1

Context

What the reader needs to know to care: "Two regions, six months of sales."

2

Contrast

The week-9 comparison — the thing that changed or differs: "Luzon leads; Visayas is climbing faster."

3

Consequence

Why it matters / what to watch: "If Visayas keeps growing twice as fast, the second half of the year is where the gap closes."

Three Beats

Setup → Tension → Resolution

The flood story, told properly.

Same three beats as the last slide under their storytelling names: context = setup, contrast = tension, consequence = resolution.

The Flood Story, In Three Beats

Each beat is one chart and one sentence

BeatThe chartThe sentence
SetupTotal spend per year, 2016–2024₱1.6 trillion over eleven years — ₱145.5B a year.
TensionContractor share, sorted descending48 of 4,841 firms hold 27.6% of it.
ResolutionThe same, with the caveat on the chartConcentration is structural, not evidence of conduct — and it is a question for the agency.
Lead With The Finding

Not with the method, and never with the dataset

You worked through the data in order: source, clean, explore, conclude. Your reader goes in the opposite direction.

Method belongs underneath

It must be there — it is what makes the claim checkable — but putting it first loses the reader before the point arrives.

# how you worked 1. downloaded DPWH export 2. cleaned 34,079 rows 3. grouped by contractor 4. found concentration # how you write 1. 48 firms hold a quarter of it 2. here is the chart 3. here is what it does not mean 4. here is the method + data # exactly reversed
The Dashboard Test

One dashboard, one question

A dashboard is one screen of a few charts that people check again and again. Before adding any chart, write the question it answers. For the lab's data: "How do the regions compare — in level and in direction?" Every panel must serve it.

  • dfthe lab's long table: 12 rows, one per month per region, columns month, region, sales
  • df.pivot_table(index=..., columns=..., values=...)a pivot table: one row per month, one column per region, sales in the cells (a wide table)
  • .reindex([...])put the rows in calendar order; pandas sorts month names alphabetically (Apr, Feb, Jan…)
  • 21.0, 14.0floats, because pivot_table averages each cell (one value each here, so the average is the value)
wide = df.pivot_table(index="month", columns="region", values="sales") wide = wide.reindex(["Jan", "Feb", "Mar", "Apr", "May", "Jun"]) print(wide) region Luzon Visayas month Jan 20.0 10.0 Feb 19.0 12.0 Mar 22.0 11.0 Apr 21.0 14.0 May 25.0 16.0 Jun 27.0 18.0

The five-second test

Show the finished dashboard to someone for five seconds. If they can't say the headline afterward, the dashboard failed — however pretty it is.

Quick Check

Tap to reveal

An office dashboard shows 14 charts — every metric the team collects. Nobody looks at it anymore. What's the root problem?

A · Needs more charts for completeness
B · It answers no particular question, so it answers nothing
C · Wrong font
D · Dashboards always fail

B — completeness is not communication.

Fourteen co-equal charts give the reader fourteen decisions about where to look — so they make none. Cut to the panels that answer one question, and the dashboard becomes worth thirty seconds again.

Break

  Five minutes

Then: assembling the argument, panel by panel.

Part B

Dashboards

Small charts, placed on purpose.

You can already draw one chart with pandas' .plot() (week 8). This part puts several small charts on one page and decides which goes where.

The Building Block

Small multiples: a grid of one-point charts

Small multiples are a grid of small charts side by side. In matplotlib the figure is the whole picture and each axes is one chart panel inside it (an odd name: it means one plotting area, not the x and y lines). plt.subplots(rows, cols) gives you a grid of axes. The rule that makes it work: each panel makes exactly one point.

tight_layout()

One call that stops titles and labels from colliding. Use it on every multi-panel figure.

fig, (a1, a2) = plt.subplots( 1, 2, figsize=(10, 3.5)) wide.plot(ax=a1, marker="o") a1.set_title("Trend by region") totals.plot.bar(ax=a2) a2.set_title("Total by region") plt.tight_layout()
  • fig, (a1, a2) = plt.subplots(1, 2, ...)make one figure with 1 row × 2 panels; call the panels a1 and a2. figsize is width, height in inches
  • wide.plot(ax=a1, marker="o")draw the pivot table from the last part on panel a1: one line per region, a dot at each month
  • totals.plot.bar(ax=a2)totals is sales summed per region (Luzon 134, Visayas 81); draw them as bars on panel a2
  • a1.set_title(...)write the text above that panel
  • What you getnothing printed: a picture with two charts side by side
What That Code Draws

Two panels: direction on the left, level on the right

One figure with two panels. Left, titled Trend by region: two lines with a dot per month from Jan to Jun, Luzon in blue rising from 20 to 27 and Visayas in orange rising from 10 to 18, with a legend. Right, titled Total by region: two bars, Luzon at 134 and Visayas at 81, with the region names written sideways.
Why Two Panels

Trend and total answer different halves

The lab's pair: a line chart shows direction (Visayas climbing fast), a bar chart shows level (Luzon still far ahead). Either alone tells half a truth.

Only the line

"Visayas is winning!" — it grew 80% (10 → 18) vs Luzon's 35% (20 → 27). True, and misleading alone.

Only the bars

"Luzon dominates, 134 vs 81." Also true, also misleading alone — Visayas went from half of Luzon's monthly sales (10 vs 20) to two-thirds (18 vs 27).

Together

"Luzon leads; Visayas is growing more than twice as fast." Now the reader knows what's actually happening.

Chart Choice

Pick the chart from the question, not from the menu

The reader's questionChart that answers itWhy
How did it change over time?Line chartThe eye follows a line left to right, like time
Which group is biggest?Bar chart, sortedBars start at zero, so their lengths compare fairly
How are the values spread out?Histogram or box plot (week 7)Shows the shape: the bulk, the tail, the outliers
Do two numbers move together?Scatter plot (week 8)One dot per record; a pattern in the cloud is a relationship
What is one headline number?A big number tile (KPI)A chart of one number is just a slower way to read it
Guiding The Eye

Readers scan; layout is your steering wheel

Left-to-right, top-to-bottom — the Z-path, the zig-zag route a reader's eye takes across a page. Whatever sits top-left gets read; whatever sits bottom-right gets skipped. And within a chart, sorting does the same job.

  • totals.sort_values(ascending=False)reorder the values; ascending=False means biggest first
  • \ at the end of a line"this line continues on the next one"
  • .plot.bar(ax=ax)draw the sorted values as bars on the panel ax
  • The bars on the righta text sketch of the chart's order, not printed output; the lab adds a third region, Mindanao (108)
totals.sort_values(ascending=False)\ .plot.bar(ax=ax) Luzon ████████████████ 134 Mindanao ████████████ 108 Visayas █████████ 81

Sorted beats alphabetical

Alphabetical order makes the reader do the ranking in their head. Sorted order is the ranking — read at a glance.

What That Code Draws

Biggest first: the order is the ranking

the bars step down from left to right, Luzon 134, Mindanao 108, Visayas 81, so the reader gets the ranking without comparing anything. This uses the lab's three-region totals and a blank fig, ax = plt.subplots(); it has no title or axis names yet, which the lab adds.

Bar chart of three regions sorted from biggest to smallest: Luzon 134, Mindanao 108, Visayas 81, with the region names written sideways under the bars.
Assembly Rules

Placing the panels

Dashboards You Can Click

Big numbers, filters and interactivity

Part C

The Craft

Titles, ink and color — where good charts are actually won.

You now have the panels and their order. This part is the finishing: the words, the lines and the colours on each single chart.

The Highest-Leverage Edit

The title states the finding, not the variable

"Sales over time" describes the axes. "Luzon sales grew 35% over the half-year" delivers the story — and the busy reader gets it even if they read nothing else. That sentence is a headline (a finding-title): a title that states what the chart shows, with its number.

  • ax.set_title("...")ax is one chart panel; .set_title() writes the text above it
  • # not: ...a comment: Python ignores everything after # on a line
  • 35%from the chart: (27 − 20) ÷ 20 = 0.35
ax.set_title( "Luzon sales grew 35% over the half-year") # not: ax.set_title("Sales")

The verifiability rule

A finding-title must be checkable from the chart below it. If the line shows 20 → 27, "grew 35%" verifies; "will dominate next year" does not belong in a title.

What That Code Draws

A title you can check against the line

the title's claim can be checked on the line below it: Jan 20, Jun 27, and (27 − 20) ÷ 20 = 35%. This is the lab's Luzon line (ax.plot(months, luzon, marker="o")) plus the title. The y axis starts at 19, not 0, because matplotlib zooms to the data, so the rise looks steeper than it is (week 9).

Line chart of Luzon sales from Jan to Jun with a dot per month, rising from 20 to 27 with a dip in Feb and Apr, titled Luzon sales grew 35% over the half-year; the y axis runs from 19 to 27.
Titles Do The Work

A title is a sentence, not a label

"Budget vs Duration" names the axes, which the axes already do. Use the most-read text on the chart to say what the chart shows.

Test: can the title stand alone?

If someone screenshots only your title, do they learn something true? If not, it is a label.

# label — wastes the best line "Budget vs Duration" # finding — earns it "The largest projects take twice as long: 269 vs 119 median days" # and for the concentration chart "48 of 4,841 contractors hold 27.6% of flood control spending"
One Colour Per Thing, Everywhere

Consistency is what makes a dashboard readable

Everything Competes With The Data

Remove until only the message is left

Gridlines, borders, tick marks, legends, background fills and drop shadows all consume attention. None of them is the finding. Decluttering is deleting them.

The subtraction pass

After the chart is right, delete one element at a time. If the message survives, it stays deleted.

  • ax.spines[["top", "right"]]spines are the four border lines of a chart; hide the top and right ones
  • ax.grid(axis="y", alpha=0.3)faint horizontal grid lines only; alpha is see-through-ness, 0 (invisible) to 1 (solid)
  • ax.tick_params(length=0)remove the little dashes (tick marks) beside the axis numbers
  • ax.legend(frameon=False)keep the legend (the colour key) but drop the box around it
ax.spines[["top", "right"]].set_visible(False) ax.grid(axis="y", alpha=0.3) ax.tick_params(length=0) ax.legend(frameon=False) # and delete outright: # 3-D effects # background fills # redundant legends (1 series) # decimal places nobody reads # PHP 267.7B, not PHP 267,688,008,746.34
What That Code Draws

Before and after the four lines

Default pandas line chart of Luzon and Visayas sales, Jan to Jun: a full box border, tick marks on both axes, and a legend in a white box headed region.The same chart after the four decluttering lines: no top or right border, faint horizontal gridlines, no tick marks, and a legend with no box and no heading.
Round For Humans

Precision the reader cannot use is noise

₱267,688,008,746.34 (Region III's total) is accurate and unreadable. ₱267.7B is accurate enough and lands instantly.

Keep full precision in the data

Round at the point of display only. Your calculation stays exact; your label stays legible.

  • df.budget.sum()add up the whole budget column: 1.6 trillion pesos, printed with every digit
  • 1e121 followed by 12 zeros (a trillion); 1e9 is a billion
  • f"PHP {total/1e12:.2f}T"an f-string: the value in { } is dropped into the text; :.2f means "2 decimal places"
  • 'PHP 1.60T'the label a human reads; the quotes mean it is now text, not a number
# in the notebook — full precision total = df.budget.sum() total 1600088782392.8298 # on the chart — human scale f"PHP {total/1e12:.2f}T" 'PHP 1.60T' # rule of thumb: 2-3 significant # figures in any label a human reads
Less Ink, More Signal

Everything on the chart competes with the data

Annotation Is A Layer, Not A Decoration

Write the conclusion onto the chart

An annotation is a short note written on the chart, next to the mark it explains. The reader should not have to locate your point among the marks.

One per chart

Two annotations usually means the chart carries two messages, which is a signal to split it.

  • ax.annotate("48 firms = ...",the note's text
  • xy=(48, 0.276)the point the arrow touches: x = 48 firms, y = 0.276 (27.6%)
  • xytext=(400, 0.45)where the text sits, away from the data
  • arrowprops=dict(...)the arrow's look: shape "->", grey colour
  • ax.axhline(0.276, ls="--", ...)a dashed horizontal reference line at 27.6%; lw is line width
ax.annotate( "48 firms = 27.6% of value", xy=(48, 0.276), xytext=(400, 0.45), arrowprops=dict(arrowstyle="->", color="#555")) # and a reference line for the threshold ax.axhline(0.276, ls="--", c="#999", lw=1)
Three Layers, In Order

Data, then emphasis, then explanation

Why Grey-Plus-One-Colour Works

The eye finds some things before you read

A pre-attentive attribute is a visual property the eye picks out in a split second, before any conscious reading: colour, size, position, length, shape. The emphasis layer uses exactly this.

One red apple in a bowl of green ones: you do not search for it, it finds you. Make your finding the red apple and everything else the green ones.

Count the 3s

65173660142313703557
37934077201250735105
16818148078587695220
05916888934153173373
27920935253003725755

Now count them again

65173660142313703557
37934077201250735105
16818148078587695220
05916888934153173373
27920935253003725755

Same digits, same 15 threes

The first block needs a careful scan, line by line. In the second, colour does the searching for you. That is what one highlighted bar does on a grey chart.

Color Is Meaning

One color per thing, everywhere

If Visayas is amber in panel one, it's amber in every panel. Consistent color lets the reader learn the mapping once — then read the rest of the dashboard for free.

Inconsistent color

Every panel re-teaches the legend; the reader spends their attention re-decoding instead of understanding.

Decorative color

Rainbow bars where color encodes nothing teach the reader to ignore color — right before you need it to mean something.

And check accessibility

Red–green pairs fail for many readers. Prefer palettes that survive grayscale, or double-encode with position and labels.

Not Everyone Sees Your Chart

Roughly 1 in 12 men cannot distinguish red from green

In a class of 40 that is two or three people. On a public dashboard it is thousands. Red-green is both the most common deficiency and the most common default palette.

Words in the code

sns
seaborn, a plotting library built on matplotlib
palette
the list of colours charts use, in order
cmap
colour map: colours for a range of numbers
sequential
light to dark for low to high
diverging
two colours meeting at a neutral middle

Encode twice

Colour plus shape, or colour plus direct labels. Then the chart survives greyscale printing too — which is the same problem in disguise.

# safe palettes sns.set_palette("colorblind") cmap="viridis" # sequential cmap="RdBu_r" # diverging, safe # NOT safe # red/green categorical pairs # "jet" / rainbow for magnitude # test it: print in greyscale. # if the message dies, so does it # for ~8% of your readers
Contrast And Size

The chart that works on your laptop and fails in the room

Default matplotlib text is sized for a notebook viewed at arm’s length. Projected, it is unreadable from the third row.

One line fixes most of it

context="talk" scales every text element at once — titles, labels, ticks and legend together.

sns.set_theme(context="talk") # paper < notebook < talk < poster # and check the things that stay small ax.tick_params(labelsize=12) # thin grey gridlines vanish on a # projector — darken or drop them ax.grid(alpha=0.4, linewidth=0.8)
Week 9 Rules, Amplified

A dashboard multiplies whatever it carries

"A misleading chart in a notebook fools one analyst. The same chart on a dashboard, refreshed daily and trusted by default, fools an organization on a schedule."
Zero-based bars, full windows, rates over raw counts — now with compound interest
Dashboards Multiply Whatever They Carry

Including the mistakes

A dashboard is not a report with more charts. It is a thing someone checks repeatedly, without you present to explain it.

The elapsed-time trap, automated

Spend per year drops from ₱369.3B (2024) to ₱196.2B (2025), and 2026 has one project. A year still being filled in when the file was exported looks like a collapse: week 8’s progress-vs-year lesson (less time, not less work). In a report you catch it once. On a dashboard it redraws every day, and every viewer draws the same wrong conclusion.

# guard the incomplete period latest = df.year.max() if df[df.year == latest].shape[0] < expected: note = f"{latest} is partial" ax.axvspan(..., alpha=.1, color="grey") ax.annotate(note, ...) # 2024: 5,553 projects # 2025: 3,475 <- still being filled in? # 2026: 1 <- exclude
  • latest = df.year.max()the most recent year in the data
  • df[df.year == latest].shape[0]keep only that year's rows, then count them (shape[0] = number of rows)
  • if ... < expected:if fewer rows than a full year normally has (you choose expected, e.g. last year's count)…
  • ax.axvspan(...), ax.annotate(...)…shade that stretch of the x-axis light grey and write the note on it; ... marks arguments left out
One Dashboard, One Question

If it answers three, it is three dashboards

The Layout Grid

Big number, then trend, then breakdown, then detail

Four zones, in reading order: the KPI tile (the big number), the trend, the breakdown, the detail table. It works because it moves from "what" to "why" the way a reader’s questions arrive.

Detail last, and optional

The table nobody reads should be at the bottom, where nobody has to scroll past it to reach the point.

┌────────────┬───────────────────┐ │ PHP 1.60T │ spend per year │ │ 11 years │ (line, 2016-24) │ ├────────────┴───────────────────┤ │ by region (bar, sorted) │ ├────────────────────────────────┤ │ top contractors (table) │ └────────────────────────────────┘ # headline -> trend -> breakdown -> detail
Presenting It

The slide is not the talk

If every word is on the slide, the audience reads ahead and stops listening. The chart carries the evidence; you carry the argument.

Say the title, then be quiet

Give people four or five seconds to read a new chart before you talk over it. It feels endless to you and is barely enough for them.

# per chart, out loud: 1. "This is spend per year, 2016 to 2024." (what) 2. ...pause. let them look. 3. "It more than triples from 2021 to 2024." (the point) 4. "Which raises the question of..." (the bridge) # what / pause / point / bridge
When You Do Not Know

The three words that protect your credibility

You will be asked something the data cannot answer. Guessing is the only response that actually costs you.

"I don't know" is a complete answer

Followed by what it would take to find out. That answer demonstrates you know where your evidence ends — which is the whole skill.

# the three legitimate answers "I don't know — the dataset has no outcome field, so I can't say whether flooding decreased." "That's outside what I measured. I looked at awarded value, not disbursement (money paid out)." "I'd have to check — let me come back to you." # never: "probably", "I think maybe"
The Six Failures

Everything that goes wrong, and it is always one of these

FailureLooks likeFix
No audience chosenMethod-heavy and also shallowName the reader before writing
Everything includedFifteen charts, no argumentThree beats, three charts
Labels as titles"Budget vs Duration"Make the title a sentence with the finding
Buried leadConclusion on the last slideLead with it; method underneath
Unmarked caveatsA partial year drawn as a collapsePut the limit on the chart, not in the notes
Talking over the chartAudience reading while you speakTitle, pause, point, bridge
Your Turn · 8 min

Build the three-beat story

1 · Choose your finding

Any verified result from weeks 4–9. It must be one you can state in a sentence with a number in it.

2 · Write three titles

Setup, tension, resolution — as full sentences. Do this before making any chart.

3 · Make exactly three charts

One per title. If you need a fourth, your story has two points; pick one.

4 · Hand it over silently

Give it to someone without speaking. Whatever you feel the urge to explain out loud is what the chart is failing to say.

Quick Check

Tap to reveal

Your dashboard shows spend per year and the latest bar is half the previous one. The export was taken mid-year. What is the right fix?

A · Leave it — the number is accurate
B · Exclude or visibly mark the incomplete period, and label it
C · Scale the partial year up to a projected annual figure
D · Remove the year axis so the drop is less obvious
B. The bar is accurate and the chart is still misleading, because a reader assumes comparable periods. C invents data. D hides the problem instead of fixing it. Mark it (greyed, hatched, annotated) or drop it — and on a dashboard this matters more than in a report, because it will redraw wrongly every day without you there.
Quick Check

Tap to reveal

Which chart title is doing its job?

A · "Budget vs Duration"
B · "The largest projects take twice as long: 269 vs 119 median days"
C · "Figure 3: Analysis of Project Data"
D · "Duration (days) by Budget (PHP)"
B. It is a sentence, it contains the finding, it carries the numbers, and it survives being screenshotted alone. A and D label the axes, which the axes already do. C tells the reader nothing at all. The title is the most-read text on any chart — spend it on the message.
Quick Check

Tap to reveal

Same chart, four titles. Which one earns its place on a dashboard?

A · "Sales"
B · "Figure 3"
C · "Luzon leads, but Visayas grew 80% since January to Luzon's 35%"
D · "Regional Sales Performance Overview Dashboard Chart"

C — it does the reader's work for them.

A and B label; D decorates; C reports — and a reader can verify it against the lines below. The title is the one part of a chart everyone reads. Spend it on the finding.

This Week's Lab

Build a narrative dashboard

Six months of sales, two regions. You'll pivot the data, pair a trend panel with a totals panel, sort for the eye, and write finding-titles a stranger could verify. ~45 minutes.

You'll practise

pivot_table, plt.subplots(1, 2), sorted plot.bar(), tight_layout(), and finding-titles.

Stretch, if you're quick

A full 2×2 dashboard, and one consistent color per region across every panel.

Recap

Four things to carry out

Glossary

This week's words, one line each

Words for today

audience
the particular reader you are writing for; decide first
narrative arc
the story's order: context → contrast → consequence
headline (finding-title)
a title that states the finding, with its number
annotation
a note on the chart, next to the mark it explains
chart choice
picking the chart from the question: line for time, bars to compare
pre-attentive attribute
colour, size, position, length: seen before reading
decluttering
deleting everything that does not help the message
colour meaning
one colour per thing, the same everywhere; grey for the rest
dashboard
one screen of a few charts that answers one question
KPI (key performance indicator)
the big headline number a dashboard tracks
small multiples
a grid of small charts, one point each
filter
a control that narrows what every panel shows
interactivity
charts that respond: hover, click, zoom
pivot table
rows = one column's values, columns = another's
figure & axes
the whole picture & one chart panel inside it
legend
the key saying which colour is which

Also new today

curse of knowledge
forgetting the reader does not know what you know
n
how many things were counted
long / wide table
one row per month-and-region / one row per month
Z-path / F-pattern
how eyes scan a page: from the top-left
data-ink
the ink that actually draws numbers
spines
the four border lines around a chart
alpha
see-through-ness, 0 (invisible) to 1 (solid)
f-string
f"...{x}...": text with a value dropped in
1e12
scientific notation: 1 followed by 12 zeros
palette / colour map
the colours used for groups / for a range of numbers
sequential / diverging
light-to-dark / two colours meeting in the middle
partial period
a year or month not yet over; mark it or drop it
Readings

Before next week

Next Week

Ethics & Privacy

In data and analytics — consent, anonymity, bias and the law: what you may do with what you can do.

DS 227 · Knowledge Discovery in Data