MATH U113 · First-time lesson · R lab support

Describing data, taught from zero

Devore (9th ed., Metric) §1.1–1.4 — populations and samples, pictures of data, centre, and spread. Budget 60–90 minutes, in two sittings if you like.

How to use this page

Scope update (handout, 01 Aug): Chapter 1 is not in the lecture plan — no lecture covers it, and it isn't listed for the midsem or compre. Where it lives instead: the Week 3 R-lab tutorial ("descriptive measures and data visualization"), which feeds the 10% Lab Exam in R. So read this page once before Week 3, then let the R lab be where you exercise it — don't spend exam-prep hours here. First time through → this page, attempting every green check box; revising → the notes page. And relax into it: most of Chapter 1 genuinely is 11th-class statistics wearing engineering clothes.

1 · Why engineers need statistics at all §1.1

Pick up a resistor marked 1 kΩ and measure it precisely: 987 Ω. Another from the same reel: 1,012 Ω. A third: 996 Ω. Nothing is broken — variability is a fact of every manufacturing process and every measurement. Identical inputs, identical machine, different outputs. Statistics is the discipline of making honest statements in the presence of that variability, and that's why every engineering degree teaches it.

Two words carry the whole course, so let's get them exactly right:

A population is the complete collection of objects you actually care about — every resistor the factory will make this year. A sample is the subset you actually measured — the 50 you pulled off the line this morning. Nearly always, the population is what you want to know about and the sample is all you can afford to look at. (Measuring everything — a census — is usually impossible or absurd: some tests destroy the item.)

That one distinction splits the whole subject in two:

So Chapter 1 is deliberately modest: no leaps, no probability — just learning to look at data properly. Don't let the modesty fool you; the vocabulary defined here (especially x̄ and s) gets used in every single later chapter.

Check yourself: a lab tests 40 batteries to destruction to see how long they last. Population? Sample? Why couldn't this be a census?

The population is all batteries of this type (including ones not yet made); the sample is the 40 tested. A census is impossible for the best possible reason: the test destroys the battery — measure the whole population and you have no product left to sell.

2 · First, always: a picture §1.2

A list of 60 numbers is unreadable. The first professional reflex with any new dataset is to draw it, and §1.2 gives you three hand tools:

A stem-and-leaf display — split each value into a stem (leading digits) and leaf (last digit), and stack the leaves on their stems. Nine bond-strength readings — 62, 65, 58, 71, 64, 68, 55, 66, 73 — become:

5 | 5 8
6 | 2 4 5 6 8
7 | 1 3

Thirty seconds of work, and you can suddenly see the data: clustered in the 60s, no stragglers, roughly symmetric. Best for small datasets (say, under 50 values) — and unlike a histogram it loses nothing: every original value is still readable.

A dotplot — one dot per value along a number line, stacking repeats. Same job, better when values repeat or when comparing two small groups.

A histogram — the workhorse for larger datasets. Chop the measurement axis into equal-width intervals ("classes"), count how many values land in each, and draw a bar of that height over each class. The bars' shape is the data's shape.

Here's the part that's easiest to learn by touch: the picture depends on how many classes you use. Below are 60 real-ish component lifetimes. Slide the control and watch the same data tell different stories:

100 260 420 lifetime (hours)

What you should notice: with 3–4 classes the data looks like a featureless lump (too much smoothing); with 20+ it's jittery noise (too little). Around 8–12 classes the true story appears — a peak near 150–200 hours and a long tail of survivors stretching right. A reasonable default the book suggests: number of classes ≈ √n (here √60 ≈ 8). There's no single "correct" histogram — which is itself worth knowing before you trust anyone else's.

That long right tail has a name you'll use constantly. A histogram is symmetric if the two halves mirror; positively (right-) skewed if the right tail stretches further; negatively (left-) skewed if the left one does. It's unimodal with one peak, bimodal with two (bimodal usually whispers: two different populations got mixed — two machines, two operators).

Check yourself: the lifetime data above — skewed which way? And (thinking ahead) will its mean sit above or below its median?

Right-skewed: the tail of long-lived components stretches toward high values. The mean gets pulled toward the tail — those few 300–400 hour survivors drag the average up — so mean > median. (Section 3 makes this precise; it's a favourite conceptual exam question.)

One video, if you want one

StatQuest: Histograms, Clearly Explained (~4 min). Watch it once, after playing with the widget above. Near the end he starts using histograms to talk about probability distributions — that's Chapter 3's business, stop there.

One refinement before moving on, because it's a classic mark-loser: all of this assumed equal-width classes. If the classes have unequal widths (books and instructors do this for data with long tails), bar height must be density = relative frequency ÷ class width, so that area, not height, represents proportion. Otherwise a wide class looks artificially tall just for being wide. The notes page works a full density-histogram example — worth a careful read, because "draw the histogram" on an exam quietly becomes this whenever widths differ.

3 · The centre of the data §1.3

Now we compress the picture into numbers. First: where is the data centred? Two honest answers, and the gap between them is the interesting part.

The sample mean is the familiar average. For sample values x₁, x₂, …, xn:

x̄ = Σxin = x₁ + x₂ + ⋯ + xnn

(Notation matters here: x̄ — "x-bar" — is the sample mean; the population mean gets the Greek letter μ. Same recipe, different data, and the whole of inference later is about using x̄ to guess μ — so the book is strict about the symbols, and exams inherit the strictness.)

Physically, the mean is the balance point: print the dotplot on cardboard and it balances on a knife-edge at x̄. That image explains the mean's one weakness — a single far-out point has enormous leverage, like a child sitting far along a see-saw.

The sample median x̃ ignores leverage entirely: sort the data; the median is the middle value (odd n) or the average of the middle two (even n). Watch them disagree. Five commute times, in minutes:

22, 25, 26, 28, 94 → x̃ = 26, x̄ = 1955 = 39

One horror commute dragged the mean to 39 — a number that describes none of your days. The median calmly reports 26. Neither is "right": the mean answers "what's my total time over many days, per day?" (leverage is the point — every minute counts), the median answers "what's a typical day?" (leverage is a lie). Choosing the honest summary for the question asked is the actual skill §1.3 teaches, and skew is the tell: mean pulled toward the tail, median not. Income statistics quote medians for exactly this reason.

Between the extremes sits the trimmed mean: delete the smallest and largest few percent, average the rest. A 10% trimmed mean of 10 values deletes the top and bottom value and averages the middle 8 — robust like the median, but still using most of the data like the mean. (The notes page's Worked Example 1 computes all three on one dataset — read it after this section.)

Finally, generalising the median: the median splits the sorted data 50/50, and quartiles split it 25/25/25/25. Devore's version of quartiles is called fourths: the lower fourth is the median of the smaller half of the data, the upper fourth the median of the larger half. Percentiles push the same idea to any split. These become the skeleton of the boxplot in section 5.

Check yourself: for the sample 3, 5, 7, 9, 100 — median, mean, and which describes a "typical value" better?

Sorted already. Median = 7 (third of five). Mean = 124/5 = 24.8. The median, by a mile: four of the five values sit between 3 and 9, and 24.8 describes none of them — the 100 is doing all the talking. Outlier-resistance is the median's whole job.

4 · The spread of the data §1.4

Two datasets can share a centre and be utterly different: {49, 50, 51} and {0, 50, 100} both have mean 50. A centre without a spread is half a description — for an engineer often the less important half. (A bridge designed for the average load fails; it's the spread that kills.)

The crudest spread measure is the range, largest minus smallest — two data points' opinion, everyone else ignored. We can do better by asking: how far is each point from the mean, typically? The deviations xi − x̄ capture this, but averaging them raw gives exactly 0 every time — negatives cancel positives (the balance point again!). So square each deviation first, making everything positive and punishing big deviations hardest. That gives the sample variance:

s2 = Σ(xi − x̄)2n − 1

and its square root, the sample standard deviation s, which undoes the squaring so the answer is back in the original units (hours, ohms, minutes — not hours²). Small worked pass, on {2, 4, 6, 8}: mean x̄ = 5; deviations −3, −1, 1, 3; squares 9, 1, 1, 9, summing to 20; so s2 = 20/3 ≈ 6.67 and s ≈ 2.58. Read s as "a typical distance from the mean" and you'll rarely be misled.

Now the question everyone asks — why divide by n − 1 and not n? Two ways to see it, both worth having:

School formulas divide by n — that's the population variance σ2, correct when you truly have every member. With a sample (almost always), it's n − 1. This single discrepancy with 11th-class habit is the most reliable mark-loser in the chapter.

Check yourself: compute s² and s for the sample 1, 2, 3, 4, 5.

x̄ = 3. Deviations: −2, −1, 0, 1, 2 → squares 4, 1, 0, 1, 4, sum 10. s2 = 10/(5−1) = 2.5, s = √2.5 ≈ 1.58. (Divided by 5? That's the n−1 trap — see above, then never again.)

5 · The boxplot: five numbers, one picture §1.4

The last tool marries pictures and numbers. Take the median, the two fourths (section 3), and the extremes — five numbers — and draw: a box from lower fourth to upper fourth, a line at the median, whiskers out to the most extreme non-suspicious points. The box holds the middle 50% of the data; its width is the fourth spread fs = upper fourth − lower fourth, a spread measure that (unlike s) shrugs off outliers.

And the boxplot comes with a built-in outlier detector — the first objective rule you've met for a word people usually use by vibes:

outlier: farther than 1.5 fs from the nearest end of the box · extreme outlier: farther than 3 fs

Quick pass on the sample 4, 5, 6, 7, 7, 8, 9, 10: median (7+7)/2 = 7; lower half {4, 5, 6, 7} has median 5.5 (lower fourth), upper half {7, 8, 9, 10} has median 8.5 (upper fourth); so fs = 3 and the outlier fences sit at 5.5 − 4.5 = 1 and 8.5 + 4.5 = 13. Every value is inside — no outliers, whiskers run to 4 and 10.

Why you'll actually use this: side-by-side boxplots are the fastest honest comparison of two processes ("supplier A vs. supplier B") — centres, spreads, skew and outliers all visible in one glance. The notes page's Worked Example 3 runs the full exam version, outlier verdicts and all.

Check yourself: same box (fourths 5.5 and 8.5) — a new reading of 14 arrives. Outlier? Extreme?

fs = 3, so the outlier fence is 8.5 + 1.5(3) = 13 and the extreme fence is 8.5 + 3(3) = 17.5. Since 13 < 14 < 17.5: a mild outlier — plotted as its own point beyond the whisker, not swept under the box.

6 · You're ready — what to do next

That's all of Chapter 1: population vs. sample, three pictures, two-and-a-half centres, spread by s and by fs, and the boxplot. To turn it into marks:

Still stuck on something? Ask an AI well

For interactive back-and-forth, a chat AI is a genuinely good first-time tool — but pin it to your syllabus so it doesn't wander. Two prompts that work:

"I'm studying Devore 9th ed. §1.4. Explain why sample variance divides by n−1, assuming I've just learned what a sample is. Then give me one 6-value dataset, let me compute s² myself, and check my work step by step."

"Using Devore §1.2–1.3 ideas only: describe three histograms in words (one symmetric, one right-skewed, one bimodal) and quiz me — for each, where does the mean sit relative to the median, and what real situation could produce it? Reveal answers only after I commit."

One caution: AI answers can contain confident arithmetic errors — cross-check any computed number by hand or against the textbook's answers in the back.