MATH U113 · Probability & Statistics · R lab support
Overview & Descriptive Statistics
Devore (9th ed., Metric) §1.1–1.4 — populations and samples, pictures of data, measures of location, measures of variability — plus the Week 3 lab sheet in R (with its summary table corrected). The compressed revision map.
Scope update (handout, 01 Aug): Chapter 1 is not in the lecture plan and not listed for the midsem or compre. It supports the Week 3 R-lab tutorial and the 10% Lab Exam in R — so treat this page as lab prep, not exam prep, and budget your hours accordingly. Meeting the material for the first time? Start with the lesson — same sections, taught slowly with the interactive histogram. And a grounding fact: most of this chapter (mean, median, variance) is 11th-class statistics for everyone; the parts that are new — dividing by n−1, fourths and boxplots — are new to the whole hall, JEE background or not.
1.1 · Populations, samples, and two kinds of statistics §1.1
A population is every object of interest; a sample is the subset actually measured; a census (measuring all) is rare — costly, slow, sometimes destructive. A variable is any characteristic whose value changes across objects; data is univariate / bivariate / multivariate by how many variables per object.
- Descriptive statistics — summarise the sample you have: pictures (§1.2) and summary numbers (§1.3–1.4). This chapter.
- Inferential statistics — generalise from sample to population, with probability (Ch 2–4) quantifying the risk in the leap. The rest of the course.
- Sampling matters: conclusions deserve trust only if the sample was sensibly drawn — the gold standard is the simple random sample (every subset equally likely); stratified sampling (sample within subgroups) is the common refinement.
1.2 · Pictures of data §1.2
| Display | Best for | Rules that earn marks |
|---|---|---|
| Stem-and-leaf | Small datasets (≲ 50); keeps every value readable | Stems = leading digit(s), leaves = last digit, one leaf per value, leaves ordered; always state the units ("stem 5, leaf 8 = 58 MPa") |
| Dotplot | Small datasets with repeats; quick comparisons | One dot per value on a number line; stack repeats |
| Histogram | Larger datasets; the shape workhorse | Equal-width classes: bar height = frequency (or relative frequency). A value on a boundary goes in the upper class ("left-inclusive" [a, b) convention) |
Shape vocabulary (asked directly in exams): symmetric / positively (right-) skewed (tail to the right) / negatively (left-) skewed; unimodal / bimodal (bimodal usually means two mixed populations). Under right skew, mean > median — the tail pulls the mean, not the median.
The unequal-width rule — the section's classic mark-loser. If class widths differ, bar height must be the density:
Plot raw frequencies on unequal classes and a wide class looks tall merely for being wide. (The "area = proportion" idea returns as the definition of a continuous distribution in Chapter 4 — this little rule is quietly foundational.)
Worked example 1Density histogram with unequal class widths
Leakage currents (mA) for 40 components, grouped as the instrument reports them:
| Class (mA) | [0, 5) | [5, 10) | [10, 20) | [20, 40) |
|---|---|---|---|---|
| Frequency | 10 | 15 | 10 | 5 |
- Relative frequencies (divide by n = 40): 0.25, 0.375, 0.25, 0.125. (Check: they sum to 1.)
- Widths: 5, 5, 10, 20 — unequal, so the density scale is mandatory.
- Densities = rel. freq ÷ width: 0.25/5 = 0.050, 0.375/5 = 0.075, 0.25/10 = 0.025, 0.125/20 = 0.00625.
- Check by area: 0.050·5 + 0.075·5 + 0.025·10 + 0.00625·20 = 0.25 + 0.375 + 0.25 + 0.125 = 1 ✓. Drawn to scale:
Notice [0, 5) and [10, 20) hold the same 10 observations each, but the [10, 20) bar is half as tall — its observations are spread over double the width. That's the honesty the density scale buys.
1.3 · Measures of location §1.3
| Measure | Recipe | Character |
|---|---|---|
| Sample mean x̄ | Σxi / n | Balance point; uses every value; sensitive to outliers |
| Sample median x̃ | Sort; middle value (odd n) or mean of middle two (even n) | Immune to outliers; ignores magnitudes in the tails |
| Trimmed mean | Delete smallest and largest k% of values, average the rest | The compromise; 100·k chosen in advance |
| Fourths (Devore's quartiles) | Lower fourth = median of the smaller half; upper fourth = median of the larger half (odd n: include the median in both halves) | Skeleton of the boxplot (§1.4) |
Symbol discipline: x̄ and s describe samples; μ and σ describe populations. Exams check this — it is the notational spine of the whole course.
Worked example 2Mean vs. median vs. trimmed mean on one dataset
Lifetimes (hours, ×100) of 10 battery packs: 5.6, 5.1, 6.2, 6.0, 5.8, 6.5, 5.8, 5.5, 5.2, 7.3.
- Mean: Σxi = 59.0, so x̄ = 59.0/10 = 5.90.
- Median: sort → 5.1, 5.2, 5.5, 5.6, 5.8, 5.8, 6.0, 6.2, 6.5, 7.3. Even n = 10: average the 5th and 6th values: x̃ = (5.8 + 5.8)/2 = 5.80.
- 10% trimmed mean: 10% of 10 is one value from each end — delete 5.1 and 7.3. Remaining sum 59.0 − 5.1 − 7.3 = 46.6 over 8 values: x̄tr(10) = 46.6/8 = 5.825.
- Interpret (exams ask): x̄ = 5.90 > x̃ = 5.80 — the slight right skew (that 7.3 straggler) pulls the mean up; the trimmed mean (5.825) lands between the two, as it should.
1.4 · Measures of variability §1.4
The range (max − min) uses two values and wastes the rest. The serious measures build on deviations from the mean, xi − x̄ (which sum to exactly 0 — hence the squaring):
Divide by n − 1, not n: only n − 1 deviations are free (they must sum to 0), and because x̄ sits closer to the sample than μ does, raw squared deviations run a little small — the smaller divisor compensates. The lesson unpacks both arguments slowly.
Unpack the shortcut formula
Expand the square: Σ(xi − x̄)2 = Σxi2 − 2x̄Σxi + nx̄2. Substitute x̄ = Σxi/n: the last two terms become −2(Σxi)2/n + (Σxi)2/n, leaving Σxi2 − (Σxi)2/n.
Quick two-way demonstration on 4, 5, 6, 7, 7, 8, 9, 10 (n = 8): by deviations, x̄ = 56/8 = 7 and the squared deviations 9, 4, 1, 0, 0, 1, 4, 9 sum to 28; by shortcut, Σxi2 = 420 so Sxx = 420 − 56²/8 = 420 − 392 = 28. Either way s2 = 28/7 = 4, s = 2. Use the shortcut when a calculator gives you Σx and Σx2; use deviations when x̄ is a round number.
Two properties worth quoting: adding a constant to every value leaves s unchanged (spread doesn't move when the whole dataset shifts); multiplying every value by c multiplies s by |c| (and s2 by c2). Handy for unit conversions — and an easy exam mark.
Boxplots and the outlier rules
The fourth spread fs = upper fourth − lower fourth is the box's width — a spread measure that ignores the tails entirely. The rules:
- Box from lower fourth to upper fourth, line at the median.
- Outlier: farther than 1.5 fs from the nearest fourth. Extreme outlier: farther than 3 fs.
- Whiskers run to the most extreme values that are not outliers; outliers are plotted individually (extreme ones with a different symbol).
Worked example 3Boxplot with an outlier check
Contaminant concentration (ppm) in 11 rinse samples, already sorted: 2.8, 3.1, 3.3, 3.5, 3.7, 3.8, 4.0, 4.2, 4.4, 4.6, 6.0.
- Median: n = 11 (odd) → 6th value: x̃ = 3.8.
- Fourths: odd n, so the median joins both halves. Smaller half {2.8, 3.1, 3.3, 3.5, 3.7, 3.8} → lower fourth (3.3 + 3.5)/2 = 3.4. Larger half {3.8, 4.0, 4.2, 4.4, 4.6, 6.0} → upper fourth (4.2 + 4.4)/2 = 4.3.
- Fourth spread: fs = 4.3 − 3.4 = 0.9, so 1.5fs = 1.35 and 3fs = 2.7.
- Outlier fences: below 3.4 − 1.35 = 2.05, above 4.3 + 1.35 = 5.65; extreme fences at 3.4 − 2.7 = 0.7 and 4.3 + 2.7 = 7.0.
- Verdicts: 6.0 > 5.65 but < 7.0 → a mild outlier. Nothing lies below 2.05. Whiskers therefore run to 2.8 and to 4.6 (the most extreme non-outliers) — not to 6.0.
Exam phrasing tip: don't just draw — state the fence arithmetic (step 4) and the verdict (step 5). The numbers are the marks; the picture is the garnish.
1.5 · The same in R — the Week 3 lab sheet Lab 1
The Week 3 sheet ("Descriptive Statistics and Distribution Plots in R", 17 Aug) runs everything above on one 15-value dataset. Each R call is a Devore idea you already have. The only genuinely new content is skewness and kurtosis, which Devore Chapter 1 never computes. That makes them lab-exam material only, never midsem or compre.
| R | Devore idea | Watch out |
|---|---|---|
| mean(x), median(x) | x̄, x̃ §1.3 | mean(x, trim = 0.1) is the 10% trimmed mean |
| var(x), sd(x) | s2, s §1.4 | Divides by n − 1, the same as Devore |
| fivenum(x), boxplot(x) | Min, fourths, median, max; boxplot with the 1.5fs outlier rule | These use Devore's fourths exactly |
| quantile(x) | "Quartiles" | Default recipe differs from Devore's fourths (see trap 2 below) |
| table(x) | Frequency table §1.2 | The sheet's get_modes() picks the most frequent value(s) out of this table |
| hist(x, probability = TRUE) | Density-scale histogram (Worked example 1) | Bar areas sum to 1. breaks = 8 is only a suggestion, and R rounds the bins to "pretty" edges |
| moments::skewness, kurtosis | Not in Devore: shape as a number | Two packages, two conventions (see trap 1) |
Skewness and kurtosis in one breath. Take the average cubed deviation m3 = Σ(xi − x̄)3/n and scale it by the spread. Cubing keeps the sign, so a long right tail makes the result positive. That is skewness, m3/m23/2, where m2 is the average squared deviation (divisor n). Fourth powers make big deviations count heavily, which gives kurtosis, m4/m22. A normal curve scores 3, and heavier tails score more. "Excess kurtosis" is simply that number minus 3.
Worked example 4The lab dataset, by hand and in R (with the sheet's table corrected)
Data: 12, 15, 9, 10, 18, 20, 22, 14, 15, 15, 19, 21, 14, 16, 17 (n = 15).
- Mean: Σxi = 237, so x̄ = 237/15 = 15.8.
- Median and mode: sorted, the list is 9, 10, 12, 14, 14, 15, 15, 15, 16, 17, 18, 19, 20, 21, 22. The 8th value gives x̃ = 15. The value 15 occurs three times, more than any other, so the mode is 15.
- Variance by the shortcut: Σxi2 = 3947, so Sxx = 3947 − 237²/15 = 3947 − 3744.6 = 202.4. Then s2 = 202.4/14 = 14.457 and s = 3.802.
- Skewness (the moments version): m2 = 202.4/15 = 13.493, and the cubed deviations average to m3 = −5.296. So skewness = −5.296/13.4931.5 = −5.296/49.57 = −0.107. That is essentially symmetric, with the faintest left lean.
- Kurtosis: m4 = 410.06, so kurtosis = 410.06/13.4932 = 2.25. That is a little under 3, meaning slightly lighter tails than a normal curve.
| Mean | Median | Mode | SD | Var | Skew | Kurt | |
|---|---|---|---|---|---|---|---|
| Sheet's Table 1 | 16.13 | 15 | 15 | 3.34 | 11.13 | 0.45 | 2.10 |
| Correct | 15.8 | 15 | 15 | 3.80 | 14.46 | −0.11 | 2.25 |
The sheet labels its table "example structure", but only the median and mode are right, and its skewness has the wrong sign. If your R output disagrees with the sheet, trust R (and this card). Your code is fine.
Unpack: mean > median, yet skewness is negative?
Mean above median suggests a right lean, but moment skewness weighs every deviation cubed. Here the low values 9 and 10 are far enough out to outweigh the upper values. On a small, nearly symmetric sample these two clues can disagree. With both this close to "symmetric", the honest verdict is roughly symmetric.
1) Two packages, two kurtoses. moments::kurtosis gives about 3 for normal data, while e1071::kurtosis (default type = 3) gives excess kurtosis, about 0. Load e1071 after moments and its functions take over. Calling library(moments) again does nothing, because the package is already attached. So the sheet's summary table, run after the e1071 block, quietly reports −1.04 instead of 2.25. The fix is to name the package explicitly: moments::kurtosis(data). 2) quantile() is not Devore's fourths. For Worked example 2's ten lifetimes, Devore's fourths are 5.5 and 6.2, but quantile() interpolates and gives 5.525 and 6.15. When Devore's answer is wanted, use fivenum(). 3) Multiple modes break the one-row summary. If get_modes() returns two values, the data.frame silently grows to two rows and repeats every other column. 4) Missing values: any NA makes mean() return NA unless you add na.rm = TRUE.
Next lab: R lab 2 · Simulating distributions, the r/d/p/q functions for Modules 2–3.
1) Dividing by n for sample variance — school habit. Sample ⟹ n − 1, every time. 2) Confusing Σxi2 (square, then sum) with (Σxi)2 (sum, then square) in the shortcut — they differ wildly, and swapping them is the classic way to get a negative variance, which is impossible: if your Sxx < 0, you've made this exact error. 3) Frequency-height bars on unequal-width classes — density scale, see Worked Example 1. 4) Reporting the mean of visibly skewed data as "the typical value" — say median, or say why not. 5) Your school textbook's quartile recipe may give slightly different numbers than Devore's fourths — in this course, use Devore's definition and name it.
Your minimal prerequisite kit for this module
- Σ notation: Σxi just means "add them all"; Σxi2 means square first, then add.
- Sorting and medians of small lists — including the even-n average-of-middle-two rule.
- Fractions ↔ decimals ↔ percentages, fluently (relative frequencies live here).
- Squares and square roots without fear, calculator allowed.
- Reading a frequency table — 9th-class material, returns in §1.2.
What to practise in Devore
| Skill | Where | How many |
|---|---|---|
| Population vs. sample, descriptive vs. inferential (concept questions) | §1.1 exercises | 3–4 |
| Stem-and-leaf and dotplot construction & reading | §1.2 exercises | 2–3 |
| Histograms — including at least one with unequal widths (density!) | §1.2 exercises | 4–5 |
| Mean, median, trimmed mean; effect of outliers | §1.3 exercises | 5–6 |
| s² both ways (definition & shortcut) | §1.4 exercises | 4–5 |
| Fourths, fs, boxplots, outlier verdicts | §1.4 exercises | 4–5 |
| R: rerun Worked example 4, then one §1.3–1.4 exercise's data through mean/median/sd/fivenum/boxplot, and check against the hand answer | Lab 1 sheet + §1.4 exercises | 2–3 |
Prefer odd-numbered exercises — Devore prints answers to selected odd ones in the back. One caution for the Metric Version: exercise numbering can differ from the US edition, so pick problems by section and skill, not by numbers copied from elsewhere.