MATH U113 · Probability & Statistics · Module 2
Discrete Random Variables & Probability Distributions
Devore (9th ed., Metric) §3.1–3.6 plus moment generating functions from Milton & Arnold R1 §3.4 — random variables, pmf & cdf, expected values, MGF, binomial, hypergeometric & negative binomial, geometric, Poisson. The compressed revision map.
First time with this material? The Module 2 lesson teaches it slowly, with the interactive binomial histogram. Grounding fact: the distribution framework is new to everyone — JEE classmates have met the binomial formula, but pmf/cdf thinking, expectation rules, and model choice are fresh for the whole hall. This is the course's biggest lecture block (L3–11, nine lectures) and squarely midsem territory. Prerequisites that must be warm: counting §2.3 and independence §2.5 from Module 1.
3.1–3.2 · Random variables, pmf, cdf §3.1–3.2
A random variable assigns a number to each outcome; discrete = listable values (counts). Capital X is the variable, lowercase x a value; "X = x" is an event. The pmf p(x) = P(X = x) is the complete value↔probability table, obeying p(x) ≥ 0 and Σp(x) = 1 (the "find the unknown constant" lever). The cdf is the running total F(x) = P(X ≤ x) — a staircase from 0 up to 1; recover the pmf from the jump sizes.
The two lookup rules that do all cdf work (integer-valued rvs):
Unpack the a−1
F(b) includes everything up to and including b. To keep a itself, subtract only what lies strictly below it — cumulative up to a−1. Writing the inequality down first ("at least 8" → X ≥ 8 → 1 − F(7)) prevents every off-by-one.
3.3 · Expected value and variance §3.3
- Mean (long-run average, balance point): E(X) = μ = Σ x p(x). For a function: E[h(X)] = Σ h(x) p(x).
- Linearity: E(aX + b) = aE(X) + b.
- Variance: V(X) = σ² = Σ(x − μ)² p(x) = E(X²) − μ² (the shortcut is usually faster); σ = √V, back in the original units.
- Rescaling: V(aX + b) = a²V(X), so σaX+b = |a|σ — the shift b never touches spread.
Worked example 1Everything from one pmf table
A lab counter tracks X = number of workstations in use at closing time:
| x | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| p(x) | 0.10 | 0.25 | 0.30 | 0.25 | 0.10 |
- Validity: all non-negative, sum = 1 ✓.
- cdf: running totals F(0) = 0.10, F(1) = 0.35, F(2) = 0.65, F(3) = 0.90, F(4) = 1. Then P(1 ≤ X ≤ 3) = F(3) − F(0) = 0.80.
- Mean: μ = 0(0.10) + 1(0.25) + 2(0.30) + 3(0.25) + 4(0.10) = 0.25 + 0.60 + 0.75 + 0.40 = 2.0. (Symmetric pmf → mean at the centre, a free sanity check.)
- Variance by shortcut: E(X²) = 0 + 1(0.25) + 4(0.30) + 9(0.25) + 16(0.10) = 0.25 + 1.20 + 2.25 + 1.60 = 5.30, so V(X) = 5.30 − 2.0² = 1.30, σ = √1.30 ≈ 1.14.
- Linear function: if power draw is W = 200X + 50 watts, then E(W) = 200(2.0) + 50 = 450 W and σW = 200(1.14) ≈ 228 W — the +50 baseline shifts the mean only.
Worked example 2The matching problem — n items shuffled, how many land in their own place? Quiz, 09 Sep
Four customers' PINs are assigned in a random order; X = number of customers who get their own PIN. Find the pmf, E(X), V(X). (Same problem in other costumes: hats returned at random, letters into envelopes, cards matched to positions.) Not binomial — the four "correct" events are not independent (three correct forces the fourth).
- Sample space: the 4! = 24 orders, equally likely.
- Count by k correct: choose which k are correct, C(4, k), and the other 4 − k must all be wrong — a derangement. Derangement counts to memorise (Devore never names them): D0 = 1, D1 = 0, D2 = 1, D3 = 2, D4 = 9.
Unpack this step: where 2 and 9 come from
D3: of the six orders of 123, only 231 and 312 fix nothing. D4: use Dn = (n − 1)(Dn−1 + Dn−2): D4 = 3(2 + 1) = 9. Check the table below sums to 24.
- pmf: p(0) = 9/24, p(1) = 4·2/24 = 8/24, p(2) = 6·1/24 = 6/24, p(3) = 4·0 = 0, p(4) = 1/24; total 24/24 ✓. ("Exactly 3 correct" is impossible.)
- Mean: E(X) = (0·9 + 1·8 + 2·6 + 4·1)/24 = 24/24 = 1.
- Variance: E(X2) = (8 + 24 + 16)/24 = 2, so V(X) = 2 − 1 = 1.
- Rescaling, as examined: points Y = 5X ⇒ E(Y) = 5; and V((a2 + a)X + 5) = (a2 + a)2 · 1 = 36 ⇒ a2 + a = 6 ⇒ a = 2 (reject −3).
Fact worth holding: for the matching problem with any n ≥ 2, E(X) = 1 and V(X) = 1. Build the table in an exam (it is the justification Module 2 gives you), but use 1 and 1 as the check. For n = 3 the pmf is 2/6, 3/6, 0, 1/6. Rebuilt from zero, with all 24 orders drawn: Quiz post-mortem, Q3.
3.3½ · The moment generating function R1 §3.4 not in Devore
The handout sources this topic from the reference book, Milton & Arnold R1 §3.4 — Devore 9e doesn't cover it at all. That makes lecture notes and tutorial sheets your primary sources; this section is the safety net. The lesson's MGF section builds the intuition first if this is your first meeting.
The moments of X are E(X), E(X²), E(X³), … — you already use the first two, since σ² = E(X²) − μ². The moment generating function packs all of them into one function of a helper variable t:
The whole point is one mechanical payoff — differentiate, then set t = 0:
Unpack why differentiating at 0 extracts moments
Differentiate the sum term-by-term in t: each term etxp(x) becomes x etxp(x) (chain rule — x is a constant here, t is the variable). Set t = 0: every e0 = 1, leaving Σ x p(x) = E(X). Each further derivative drops another factor of x.
Two sanity anchors: M(0) = Σ p(x) = 1 always (if yours isn't, the algebra broke), and the MGF identifies the distribution uniquely — if an exam hands you M(t) = (0.7 + 0.3et)10, you may declare X ~ Bin(10, 0.3) and answer with the binomial pmf: P(X = 2) = (10 choose 2)(0.3)²(0.7)⁸ ≈ 0.2335.
Worked example M1Derive the Poisson MGF, then read off mean and variance
X ~ Poisson(μ), p(x) = e−μμx/x! for x = 0, 1, 2, …
- Assemble the sum: M(t) = Σ etx e−μμx/x! = e−μ Σ (μet)x/x! — the two x-powers merged into one.
- Spot the series: Σ ux/x! = eu (the exponential series, with u = μet). So M(t) = e−μeμet = eμ(et−1). Check: M(0) = e0 = 1 ✓
- First derivative (chain rule): M′(t) = μet · M(t), so M′(0) = μ · 1 = μ — the Poisson mean, one line.
- Second derivative (product rule on μetM): M″ = μetM + μetM′; at 0: μ + μ².
- Variance: σ² = (μ + μ²) − μ² = μ — the famous "mean = variance" signature, earned rather than memorised.
Worked example M2The binomial MGF
X ~ Bin(n, p), q = 1−p. Derive M(t) and recover E(X) = np.
- M(t) = Σ etx(n choose x)pxqn−x = Σ (n choose x)(pet)xqn−x — again the etx merges with px.
- Spot the binomial theorem, Σ (n choose x)axbn−x = (a+b)n: M(t) = (q + pet)n. Check: M(0) = (q+p)n = 1 ✓
- Differentiate (chain rule): M′(t) = n(q + pet)n−1 · pet; at t = 0: n · 1n−1 · p = np ✓ — compare the two-page direct-sum derivation this replaces.
- (One more derivative the same way gives M″(0) = np + n(n−1)p², hence σ² = np − np² = npq.)
| Model | MGF | M′(0) |
|---|---|---|
| Binomial(n, p) | (q + pet)n | np |
| Poisson(μ) | eμ(et−1) | μ |
| Geometric — trials until first success (R1's convention) | pet1 − qet | 1/p |
| Geometric — failures before first success (Devore's convention) | p1 − qet | q/p |
| Negative binomial — trials until the r-th success (R1; mid-sem 2025 Q7) | (pet1 − qet)r | r/p |
| Negative binomial — failures before the r-th success (Devore) | (p1 − qet)r | rq/p |
1) Differentiate first, substitute t = 0 second — substituting first gives the useless constant 1. 2) e0 = 1, not 0 — the classic slip inside step 3 above. 3) The geometric has two conventions (table above) and this course uses both books: Milton & Arnold count trials until success (mean 1/p), Devore counts failures before it (mean q/p). They differ by exactly 1. On any geometric question, write down which variable you're using before computing. 4) M″(0) is E(X²), not the variance — subtracting [M′(0)]² is still on you.
The four models — and how to choose §3.4–3.6
Exams rarely say "use the binomial". They tell a story; you identify the structure. The decision table (the single most valuable thing on this page):
| Story structure | Model | pmf | E(X) · V(X) |
|---|---|---|---|
| Fixed n independent yes/no trials, same p; count successes | Binomial §3.4 | (n choose x) pxqn−x, q = 1−p | np · npq |
| Draw n from a pool of N without replacement, M marked; count marked | Hypergeometric §3.5 | (M ch x)(N−M ch n−x)(N ch n) | nMN · N−nN−1·nMN(1−MN) |
| Keep trying until the r-th success; count failures on the way (r = 1: geometric) | Negative binomial §3.5 | (x+r−1 ch r−1) prqx | rqp · rqp² |
| Count events in a time/space window, rate α, no fixed n | Poisson §3.6, μ = αt | e−μμxx! | μ · μ |
Bridges between the models (each is exam-quotable): sampling without replacement with a huge pool (n/N ≤ 0.05) → hypergeometric ≈ binomial. Binomial with n large, p tiny (n ≥ 50, np ≤ 5) → ≈ Poisson with μ = np. And Module 1's Worked Example 1 (4 boards from 12, 3 defective) was the hypergeometric all along: h(1; 4, 3, 12) = 28/55, with E(X) = 4 · 3/12 = 1.
Poisson's signature: mean = variance = μ. And its μ must be scaled to the window in the question — rate 2/min over 3 minutes means μ = 6, never 2.
Worked example 2Binomial in full (recognise → compute → cdf form)
Each of 10 finished units independently passes final inspection with probability 0.9. Find (a) P(exactly 8 pass), (b) P(at least 8 pass), (c) mean and SD of the number passing.
- Model check: fixed n = 10, two outcomes, independent, constant p = 0.9 → binomial, X ~ Bin(10, 0.9). (Write this line in exams — it carries marks.)
- (a) b(8; 10, 0.9) = (10 choose 8)(0.9)⁸(0.1)² = 45 × 0.43047 × 0.01 ≈ 0.194.
Unpack (0.9)⁸
Square three times: 0.9² = 0.81, 0.81² = 0.6561, 0.6561² ≈ 0.43047. Repeated squaring beats eight keystrokes and eight rounding errors.
- (b) Sum the top three terms: b(9) = 10(0.9)⁹(0.1) ≈ 0.387 and b(10) = (0.9)¹⁰ ≈ 0.349, so P(X ≥ 8) ≈ 0.194 + 0.387 + 0.349 = 0.930. In cdf language: 1 − B(7; 10, 0.9) — the form to use with the Appendix tables.
- (c) E(X) = np = 9; V(X) = npq = 10(0.9)(0.1) = 0.9; σ ≈ 0.95. Sanity: the answer to (b) says X ≥ 8 about 93% of the time — consistent with a mean of 9 and a small σ. ✓
Worked example 3Poisson process (scale the rate first)
Support calls arrive at a steady average rate α = 2 per minute, independently. For a 3-minute window, find (a) P(exactly 4 calls), (b) P(at most 2 calls), (c) P(at least one call), and the SD of the count.
- Scale the rate to the window: μ = αt = 2 × 3 = 6. (Using 2 here is the section's classic trap.)
- (a) P(X = 4) = e⁻⁶ 6⁴/4! = e⁻⁶ · 1296/24 = 54e⁻⁶ ≈ 0.134.
- (b) P(X ≤ 2) = e⁻⁶(1 + 6 + 6²/2) = 25e⁻⁶ ≈ 0.062 — only a 6% chance of a quiet window.
- (c) Complement: P(X ≥ 1) = 1 − e⁻⁶ ≈ 0.9975. And σ = √μ = √6 ≈ 2.45 (mean and variance are both μ — the Poisson signature).
1) cdf off-by-one: "at least 8" is 1 − F(7), not 1 − F(8) — write the inequality before touching the table. 2) V(aX+b) = a²V(X) — the b vanishes, the a squares; and variance is never negative (if yours is, the E(X²) − μ² order got flipped). 3) Calling "without replacement" binomial — it's hypergeometric unless n/N ≤ 0.05; cite the ratio when you approximate. 4) Poisson with the raw rate instead of μ = αt — scale to the window asked. 5) E(X²) ≠ [E(X)]² — the difference between them is exactly the variance; conflating them makes every V(X) come out 0.
The pmf ratio f(x+1)/f(x) — one line per model
Not a Devore section, but a 12-mark mid-sem question (2025 Q6, worked in full here) is built on it. Divide a pmf by itself one step earlier: the factorials and powers cancel, leaving a short ratio. Each model has its own fingerprint:
| Model | f(x+1) / f(x) | As α/(x+1) + β |
|---|---|---|
| Binomial(n, p) | (n − x) p(x + 1) q | α = (n+1)p/q, β = −p/q |
| Negative binomial, failures (Devore) | (x + r) qx + 1 | α = (r−1)q, β = q |
| Poisson(μ) | μx + 1 | α = μ, β = 0 |
How to split a ratio into that form: write the top in terms of (x+1). For the binomial, n − x = (n+1) − (x+1), so the ratio is (p/q)[(n+1)/(x+1) − 1]. For the negative binomial, x + r = (x+1) + (r−1).
The mean trick. Multiply f(x+1) = [α/(x+1) + β] f(x) by (x+1) and sum over all x ≥ 0. The left side is Σ (x+1) f(x+1) = E(X); the right side is Σ [α + β(x+1)] f(x) = α + βE(X) + β. Solve:
Checks: binomial gives (np/q)/(1/q) = np ✓; negative binomial gives rq/p ✓; Poisson gives μ ✓. (The trials-until convention, X ≥ r, has ratio xq/(x−r+1), which does not fit this form — the question's nb(r, p) is Devore's failures count.)
Your minimal prerequisite kit for this module
- Binomial coefficients (n choose x), fluent — this is Module 1 §2.3's payoff; revise there if rusty.
- Independence & the multiplication rule (Module 1 §2.5) — the engine inside the binomial pmf.
- Weighted averages (the sample mean from the descriptive-statistics page) — E(X) is the same computation with probabilities as weights.
- Powers on a calculator: px by repeated squaring; the ex button for Poisson (e ≈ 2.71828).
- Translating words to inequalities: "at least" → ≥, "at most" → ≤, "more than" → > — before any cdf lookup.
What to practise in Devore
| Skill | Where | How many |
|---|---|---|
| pmf validity, find-the-constant, cdf ↔ pmf both directions | §3.2 exercises | 4–5 |
| E(X), V(X) from tables; E[h(X)]; linear-function rules | §3.3 exercises | 5–6 |
| Binomial: recognise, compute pmf terms, use cdf tables | §3.4 exercises | 5–6 |
| Hypergeometric & negative binomial (confirmed in-syllabus by the handout) | §3.5 exercises | 3–4 |
| Moment generating functions (not in Devore — use the reference book) | Milton & Arnold (R1) §3.4 exercises + tutorial sheet | 3–4 |
| Poisson: direct, as a process (μ = αt), as binomial approximation | §3.6 exercises | 4–5 |
Prefer odd-numbered exercises (answers to selected odd ones in the back). Metric Version numbering may differ from the US edition — choose by section and skill, not by numbers copied from elsewhere.