MATH U113 · Probability & Statistics · R lab support

Overview & Descriptive Statistics

Devore (9th ed., Metric) §1.1–1.4 — populations and samples, pictures of data, measures of location, measures of variability — plus the Week 3 lab sheet in R (with its summary table corrected). The compressed revision map.

Start here

Scope update (handout, 01 Aug): Chapter 1 is not in the lecture plan and not listed for the midsem or compre. It supports the Week 3 R-lab tutorial and the 10% Lab Exam in R — so treat this page as lab prep, not exam prep, and budget your hours accordingly. Meeting the material for the first time? Start with the lesson — same sections, taught slowly with the interactive histogram. And a grounding fact: most of this chapter (mean, median, variance) is 11th-class statistics for everyone; the parts that are new — dividing by n−1, fourths and boxplots — are new to the whole hall, JEE background or not.

1.1 · Populations, samples, and two kinds of statistics §1.1

A population is every object of interest; a sample is the subset actually measured; a census (measuring all) is rare — costly, slow, sometimes destructive. A variable is any characteristic whose value changes across objects; data is univariate / bivariate / multivariate by how many variables per object.

1.2 · Pictures of data §1.2

DisplayBest forRules that earn marks
Stem-and-leafSmall datasets (≲ 50); keeps every value readableStems = leading digit(s), leaves = last digit, one leaf per value, leaves ordered; always state the units ("stem 5, leaf 8 = 58 MPa")
DotplotSmall datasets with repeats; quick comparisonsOne dot per value on a number line; stack repeats
HistogramLarger datasets; the shape workhorseEqual-width classes: bar height = frequency (or relative frequency). A value on a boundary goes in the upper class ("left-inclusive" [a, b) convention)

Shape vocabulary (asked directly in exams): symmetric / positively (right-) skewed (tail to the right) / negatively (left-) skewed; unimodal / bimodal (bimodal usually means two mixed populations). Under right skew, mean > median — the tail pulls the mean, not the median.

The unequal-width rule — the section's classic mark-loser. If class widths differ, bar height must be the density:

density = relative frequencyclass width ⟹ bar area = relative frequency, total area = 1

Plot raw frequencies on unequal classes and a wide class looks tall merely for being wide. (The "area = proportion" idea returns as the definition of a continuous distribution in Chapter 4 — this little rule is quietly foundational.)

Worked example 1Density histogram with unequal class widths

Leakage currents (mA) for 40 components, grouped as the instrument reports them:

Class (mA)[0, 5)[5, 10)[10, 20)[20, 40)
Frequency1015105
  1. Relative frequencies (divide by n = 40): 0.25, 0.375, 0.25, 0.125. (Check: they sum to 1.)
  2. Widths: 5, 5, 10, 20 — unequal, so the density scale is mandatory.
  3. Densities = rel. freq ÷ width: 0.25/5 = 0.050, 0.375/5 = 0.075, 0.25/10 = 0.025, 0.125/20 = 0.00625.
  4. Check by area: 0.050·5 + 0.075·5 + 0.025·10 + 0.00625·20 = 0.25 + 0.375 + 0.25 + 0.125 = 1 ✓. Drawn to scale:
0 5 10 20 40 leakage current (mA)

Notice [0, 5) and [10, 20) hold the same 10 observations each, but the [10, 20) bar is half as tall — its observations are spread over double the width. That's the honesty the density scale buys.

1.3 · Measures of location §1.3

MeasureRecipeCharacter
Sample mean x̄Σxi / nBalance point; uses every value; sensitive to outliers
Sample median x̃Sort; middle value (odd n) or mean of middle two (even n)Immune to outliers; ignores magnitudes in the tails
Trimmed meanDelete smallest and largest k% of values, average the restThe compromise; 100·k chosen in advance
Fourths (Devore's quartiles)Lower fourth = median of the smaller half; upper fourth = median of the larger half (odd n: include the median in both halves)Skeleton of the boxplot (§1.4)

Symbol discipline: x̄ and s describe samples; μ and σ describe populations. Exams check this — it is the notational spine of the whole course.

Worked example 2Mean vs. median vs. trimmed mean on one dataset

Lifetimes (hours, ×100) of 10 battery packs: 5.6, 5.1, 6.2, 6.0, 5.8, 6.5, 5.8, 5.5, 5.2, 7.3.

  1. Mean: Σxi = 59.0, so x̄ = 59.0/10 = 5.90.
  2. Median: sort → 5.1, 5.2, 5.5, 5.6, 5.8, 5.8, 6.0, 6.2, 6.5, 7.3. Even n = 10: average the 5th and 6th values: x̃ = (5.8 + 5.8)/2 = 5.80.
  3. 10% trimmed mean: 10% of 10 is one value from each end — delete 5.1 and 7.3. Remaining sum 59.0 − 5.1 − 7.3 = 46.6 over 8 values: x̄tr(10) = 46.6/8 = 5.825.
  4. Interpret (exams ask): x̄ = 5.90 > x̃ = 5.80 — the slight right skew (that 7.3 straggler) pulls the mean up; the trimmed mean (5.825) lands between the two, as it should.

1.4 · Measures of variability §1.4

The range (max − min) uses two values and wastes the rest. The serious measures build on deviations from the mean, xi − x̄ (which sum to exactly 0 — hence the squaring):

s2 = Σ(xi − x̄)2n − 1 = Sxxn − 1 where Sxx = Σxi2 − (Σxi)2n · s = √s2

Divide by n − 1, not n: only n − 1 deviations are free (they must sum to 0), and because x̄ sits closer to the sample than μ does, raw squared deviations run a little small — the smaller divisor compensates. The lesson unpacks both arguments slowly.

Unpack the shortcut formula

Expand the square: Σ(xi − x̄)2 = Σxi2 − 2x̄Σxi + nx̄2. Substitute x̄ = Σxi/n: the last two terms become −2(Σxi)2/n + (Σxi)2/n, leaving Σxi2 − (Σxi)2/n.

Quick two-way demonstration on 4, 5, 6, 7, 7, 8, 9, 10 (n = 8): by deviations, x̄ = 56/8 = 7 and the squared deviations 9, 4, 1, 0, 0, 1, 4, 9 sum to 28; by shortcut, Σxi2 = 420 so Sxx = 420 − 56²/8 = 420 − 392 = 28. Either way s2 = 28/7 = 4, s = 2. Use the shortcut when a calculator gives you Σx and Σx2; use deviations when x̄ is a round number.

Two properties worth quoting: adding a constant to every value leaves s unchanged (spread doesn't move when the whole dataset shifts); multiplying every value by c multiplies s by |c| (and s2 by c2). Handy for unit conversions — and an easy exam mark.

Boxplots and the outlier rules

The fourth spread fs = upper fourth − lower fourth is the box's width — a spread measure that ignores the tails entirely. The rules:

Worked example 3Boxplot with an outlier check

Contaminant concentration (ppm) in 11 rinse samples, already sorted: 2.8, 3.1, 3.3, 3.5, 3.7, 3.8, 4.0, 4.2, 4.4, 4.6, 6.0.

  1. Median: n = 11 (odd) → 6th value: x̃ = 3.8.
  2. Fourths: odd n, so the median joins both halves. Smaller half {2.8, 3.1, 3.3, 3.5, 3.7, 3.8} → lower fourth (3.3 + 3.5)/2 = 3.4. Larger half {3.8, 4.0, 4.2, 4.4, 4.6, 6.0} → upper fourth (4.2 + 4.4)/2 = 4.3.
  3. Fourth spread: fs = 4.3 − 3.4 = 0.9, so 1.5fs = 1.35 and 3fs = 2.7.
  4. Outlier fences: below 3.4 − 1.35 = 2.05, above 4.3 + 1.35 = 5.65; extreme fences at 3.4 − 2.7 = 0.7 and 4.3 + 2.7 = 7.0.
  5. Verdicts: 6.0 > 5.65 but < 7.0 → a mild outlier. Nothing lies below 2.05. Whiskers therefore run to 2.8 and to 4.6 (the most extreme non-outliers) — not to 6.0.
concentration (ppm)

Exam phrasing tip: don't just draw — state the fence arithmetic (step 4) and the verdict (step 5). The numbers are the marks; the picture is the garnish.

1.5 · The same in R — the Week 3 lab sheet Lab 1

The Week 3 sheet ("Descriptive Statistics and Distribution Plots in R", 17 Aug) runs everything above on one 15-value dataset. Each R call is a Devore idea you already have. The only genuinely new content is skewness and kurtosis, which Devore Chapter 1 never computes. That makes them lab-exam material only, never midsem or compre.

RDevore ideaWatch out
mean(x), median(x)x̄, x̃ §1.3mean(x, trim = 0.1) is the 10% trimmed mean
var(x), sd(x)s2, s §1.4Divides by n − 1, the same as Devore
fivenum(x), boxplot(x)Min, fourths, median, max; boxplot with the 1.5fs outlier ruleThese use Devore's fourths exactly
quantile(x)"Quartiles"Default recipe differs from Devore's fourths (see trap 2 below)
table(x)Frequency table §1.2The sheet's get_modes() picks the most frequent value(s) out of this table
hist(x, probability = TRUE)Density-scale histogram (Worked example 1)Bar areas sum to 1. breaks = 8 is only a suggestion, and R rounds the bins to "pretty" edges
moments::skewness, kurtosisNot in Devore: shape as a numberTwo packages, two conventions (see trap 1)

Skewness and kurtosis in one breath. Take the average cubed deviation m3 = Σ(xi − x̄)3/n and scale it by the spread. Cubing keeps the sign, so a long right tail makes the result positive. That is skewness, m3/m23/2, where m2 is the average squared deviation (divisor n). Fourth powers make big deviations count heavily, which gives kurtosis, m4/m22. A normal curve scores 3, and heavier tails score more. "Excess kurtosis" is simply that number minus 3.

Worked example 4The lab dataset, by hand and in R (with the sheet's table corrected)

Data: 12, 15, 9, 10, 18, 20, 22, 14, 15, 15, 19, 21, 14, 16, 17 (n = 15).

  1. Mean: Σxi = 237, so x̄ = 237/15 = 15.8.
  2. Median and mode: sorted, the list is 9, 10, 12, 14, 14, 15, 15, 15, 16, 17, 18, 19, 20, 21, 22. The 8th value gives x̃ = 15. The value 15 occurs three times, more than any other, so the mode is 15.
  3. Variance by the shortcut: Σxi2 = 3947, so Sxx = 3947 − 237²/15 = 3947 − 3744.6 = 202.4. Then s2 = 202.4/14 = 14.457 and s = 3.802.
  4. Skewness (the moments version): m2 = 202.4/15 = 13.493, and the cubed deviations average to m3 = −5.296. So skewness = −5.296/13.4931.5 = −5.296/49.57 = −0.107. That is essentially symmetric, with the faintest left lean.
  5. Kurtosis: m4 = 410.06, so kurtosis = 410.06/13.4932 = 2.25. That is a little under 3, meaning slightly lighter tails than a normal curve.
MeanMedianModeSDVarSkewKurt
Sheet's Table 116.1315153.3411.130.452.10
Correct15.815153.8014.46−0.112.25

The sheet labels its table "example structure", but only the median and mode are right, and its skewness has the wrong sign. If your R output disagrees with the sheet, trust R (and this card). Your code is fine.

Unpack: mean > median, yet skewness is negative?

Mean above median suggests a right lean, but moment skewness weighs every deviation cubed. Here the low values 9 and 10 are far enough out to outweigh the upper values. On a small, nearly symmetric sample these two clues can disagree. With both this close to "symmetric", the honest verdict is roughly symmetric.

Classic traps in the R lab

1) Two packages, two kurtoses. moments::kurtosis gives about 3 for normal data, while e1071::kurtosis (default type = 3) gives excess kurtosis, about 0. Load e1071 after moments and its functions take over. Calling library(moments) again does nothing, because the package is already attached. So the sheet's summary table, run after the e1071 block, quietly reports −1.04 instead of 2.25. The fix is to name the package explicitly: moments::kurtosis(data). 2) quantile() is not Devore's fourths. For Worked example 2's ten lifetimes, Devore's fourths are 5.5 and 6.2, but quantile() interpolates and gives 5.525 and 6.15. When Devore's answer is wanted, use fivenum(). 3) Multiple modes break the one-row summary. If get_modes() returns two values, the data.frame silently grows to two rows and repeats every other column. 4) Missing values: any NA makes mean() return NA unless you add na.rm = TRUE.

Next lab: R lab 2 · Simulating distributions, the r/d/p/q functions for Modules 2–3.

Classic traps in Chapter 1

1) Dividing by n for sample variance — school habit. Sample ⟹ n − 1, every time. 2) Confusing Σxi2 (square, then sum) with (Σxi)2 (sum, then square) in the shortcut — they differ wildly, and swapping them is the classic way to get a negative variance, which is impossible: if your Sxx < 0, you've made this exact error. 3) Frequency-height bars on unequal-width classes — density scale, see Worked Example 1. 4) Reporting the mean of visibly skewed data as "the typical value" — say median, or say why not. 5) Your school textbook's quartile recipe may give slightly different numbers than Devore's fourths — in this course, use Devore's definition and name it.

Your minimal prerequisite kit for this module

What to practise in Devore

SkillWhereHow many
Population vs. sample, descriptive vs. inferential (concept questions)§1.1 exercises3–4
Stem-and-leaf and dotplot construction & reading§1.2 exercises2–3
Histograms — including at least one with unequal widths (density!)§1.2 exercises4–5
Mean, median, trimmed mean; effect of outliers§1.3 exercises5–6
s² both ways (definition & shortcut)§1.4 exercises4–5
Fourths, fs, boxplots, outlier verdicts§1.4 exercises4–5
R: rerun Worked example 4, then one §1.3–1.4 exercise's data through mean/median/sd/fivenum/boxplot, and check against the hand answerLab 1 sheet + §1.4 exercises2–3

Prefer odd-numbered exercises — Devore prints answers to selected odd ones in the back. One caution for the Metric Version: exercise numbering can differ from the US edition, so pick problems by section and skill, not by numbers copied from elsewhere.