MATH U113 · Probability & Statistics · Module 3 · Part 2
MGFs, Chebyshev & Transformations
The three handout topics Devore doesn't carry: moment generating functions for continuous rvs (lectures L12–16, R1 §3.4 ideas carried over), Chebyshev's inequality (L19, R1 §4.5, no proof), transformation methods (L20, R1 §6.7). The compressed revision map.
First time with this material? The Part 2 lesson teaches it slowly, with two sliders. Grounding facts: none of this is in Devore, and none of it is on any school or entrance syllabus — the handout pulls it from the reference book, so every person in the hall meets it at the same time, from lecture notes. The calculus is bounded: one integral of e−cx, one completing-the-square, two standard limits.
Why it's worth your hours before 7 Oct: in the last three mid-sem papers, the 2025 paper spent 21 of its 60 marks here — Q1 (9 marks) is MGF recognition, Q7 (12 marks) is an MGF plus a limit. Chebyshev and transformations did not appear on the pages we have (the 2024 paper's page 2 is missing), but both are lectured before the mid-sem (L19–20), so treat them as fair game at the depth below.
1 · MGFs for continuous random variables L12–16 not in Devore
Same definition as Module 2's MGF, with Σ traded for ∫ — the usual Chapter 4 swap:
One new feature: the integral may only exist for some t (for the exponential, only t < λ). Always state that range with the formula — examiners mark it. The catalogue (memorise the right-hand three columns; it's a closed-book exam):
| Distribution | MGF M(t) | Valid for | Mean | Variance |
|---|---|---|---|---|
| Uniform on [a, b] | ebt − eat(b − a)t (and 1 at t = 0) | all t | (a+b)/2 | (b−a)²/12 |
| Exponential(λ) | λλ − t | t < λ | 1/λ | 1/λ² |
| Gamma(α, β) | (1 − βt)−α | t < 1/β | αβ | αβ² |
| Chi-squared(ν) = Gamma(ν/2, 2) | (1 − 2t)−ν/2 | t < ½ | ν | 2ν |
| Normal N(μ, σ²) | eμt + σ²t²/2 | all t | μ | σ² |
Two properties do most of the exam work. Uniqueness: an MGF identifies its distribution — match the shape, declare the family. Linear change: for Y = aX + b,
(ebt is a constant, so it leaves the expectation; what remains is the MGF of X evaluated at at instead of t.)
Worked example 1Exponential MGF, its range, and its moments
X ~ Exp(λ), f(x) = λe−λx for x ≥ 0. Derive M(t), then E(X) and V(X).
- Merge the exponentials: M(t) = ∫0∞ etxλe−λx dx = λ∫0∞ e−(λ−t)x dx.
- Where it exists: e−cx dies away as x → ∞ only if c = λ − t > 0. If t ≥ λ the integrand never shrinks and the area is infinite. So: t < λ.
- Integrate: ∫0∞ e−cx dx = [−e−cx/c]0∞ = 0 − (−1/c) = 1/c, so M(t) = λ/(λ − t), t < λ. Check M(0) = 1 ✓.
- Moments: write M = λ(λ − t)−1. M′ = λ(λ − t)−2 (chain rule: the inner derivative is −1, which cancels the power rule's −1) → M′(0) = 1/λ. M″ = 2λ(λ − t)−3 → M″(0) = 2/λ².
- Variance: 2/λ² − (1/λ)² = 1/λ² — Module 3's "mean = SD = 1/λ", now earned.
Gamma, the same way. M(t) = ∫0∞ etx xα−1e−x/βΓ(α)βα dx. Merge exponents: e−x(1/β − t) = e−x/θ with new scale θ = β/(1 − βt) (needs t < 1/β). Now the one trick of the section: since a gamma pdf integrates to 1, ∫0∞ xα−1e−x/θ dx = Γ(α)θα for any scale θ. So M(t) = Γ(α)θα/(Γ(α)βα) = (θ/β)α = (1 − βt)−α. Put α = 1, β = 1/λ and it collapses to Worked example 1 ✓; put α = ν/2, β = 2 for chi-squared.
The normal MGF — completing the square first
The derivation needs one piece of school algebra, so here it is before it's used. Completing the square rewrites tz − z²/2 as one perfect square plus a leftover constant. Expand (z − t)² = z² − 2tz + t²; halve and negate: −(z − t)²/2 = −z²/2 + tz − t²/2. That is our expression minus t²/2, so adding it back:
Worked example 2The normal MGF
Show MZ(t) = et²/2 for Z ~ N(0, 1), then MX(t) = eμt + σ²t²/2 for X ~ N(μ, σ²).
- Set up: MZ(t) = ∫ etz · 1√(2π)e−z²/2 dz = ∫ 1√(2π)etz − z²/2 dz (over the whole line).
- Complete the square (above): the exponent is −(z − t)²/2 + t²/2. The et²/2 part has no z, so it comes out of the integral.
- Recognise what's left: ∫ 1√(2π)e−(z−t)²/2 dz is the area under a N(t, 1) density — the standard bell slid sideways by t. Any pdf has area 1. So MZ(t) = et²/2, for every t.
- General normal: X = μ + σZ, so by the linear-change rule with a = σ, b = μ: MX(t) = eμtMZ(σt) = eμteσ²t²/2 = eμt + σ²t²/2.
- Bonus (optional — the handout omits the normal's mean/variance proofs): M′ = (μ + σ²t)M → μ; M″ = σ²M + (μ + σ²t)²M → σ² + μ²; variance σ² ✓.
2 · Reading a distribution off its MGF uniqueness
The exam move is backwards: you're handed an MGF and must name the distribution. Rewrite it into one of the catalogue shapes, then read the parameters off by matching term for term.
| If you see… | …rewrite as | …declare |
|---|---|---|
| e(quadratic in t, no constant) | eμt + σ²t²/2: μ = coefficient of t, σ² = 2 × coefficient of t² | N(μ, σ²) |
| (1 − ct)−k | Gamma with β = c, α = k | Gamma(k, c); if k = 1, exponential with mean c; if c = 2, χ²(2k) |
| λ/(λ − t) | divide top and bottom by λ: (1 − t/λ)−1 | Exp(λ), mean 1/λ |
| ebt × (a catalogue MGF) | linear-change rule with a = 1 | that distribution, shifted by b |
Worked example 32025 mid-sem, Q1 [9 marks]
Exam-completion time has MGF e0.5(σ²t² + 160t), and 30.85% of students take more than 85 minutes. Find the probability that a student finishes (a) within one hour; (b) in more than 65 but less than 75 minutes.
- Multiply the 0.5 through: exponent = 80t + σ²t²/2. That is the normal shape with μ = 80. By uniqueness, X ~ N(80, σ²) — say "by the uniqueness of MGFs" in the answer; it's the method mark.
- Find σ from the given fact: P(X > 85) = 1 − Φ(5/σ) = 0.3085, so Φ(5/σ) = 0.6915. The table says Φ(0.5) = 0.6915, so 5/σ = 0.5 and σ = 10.
- (a) "Within one hour" means X ≤ 60: z = (60 − 80)/10 = −2, Φ(−2) = 1 − 0.9772 = 0.0228.
- (b) z-values (65 − 80)/10 = −1.5 and (75 − 80)/10 = −0.5: Φ(−0.5) − Φ(−1.5) = 0.3085 − 0.0668 = 0.2417.
3 · Limits via MGFs appeared in the 2025 mid-sem
The handout doesn't list this by name, but the 2025 mid-sem gave it 12 marks (Q7). Learn the one idea and the one worked example below; confirm the depth with your instructor rather than hunting for more.
The idea: if the MGFs of Y1, Y2, … converge to the MGF of some distribution (for all t near 0), then the distributions themselves converge to it. Uniqueness, taken to the limit. The algebra always leans on one of two school limits:
(The first is the derivative of ex at 0, which is e0 = 1, written as a limit.)
Worked example 42025 mid-sem, Q7 [12 marks]
Independent Bernoulli(p) trials; X = number of trials needed for the r-th success. (a) pmf and MGF of X. (b) MGF of Y = 2pX, and the limiting distribution of Y as p → 0.
- pmf: X = x means trial x is a success and the first x − 1 trials hold exactly r − 1 successes: p(x) = (x−1 choose r−1) prqx−r, x = r, r+1, … This is the trials convention — Devore's nb counts failures X − r instead; say which you use.
- Set up the MGF, substituting x = r + k (k = failures): M(t) = Σk≥0 (r+k−1 choose k) prqket(r+k) = (pet)r Σk≥0 (r+k−1 choose k)(qet)k. ((r+k−1 choose k) = (r+k−1 choose r−1): choosing which k slots fail = choosing which r − 1 succeed.)
- Sum the series with the "pmf sums to 1" trick — the same trick as the gamma integral. Devore's negative binomial pmf sums to 1: Σ (r+k−1 choose k) prqk = 1, i.e. Σ (r+k−1 choose k) uk = (1 − u)−r for any 0 < u < 1 (write u for q, 1 − u for p). With u = qet:
- Sanity check: r = 1 gives Module 2's geometric (trials) MGF ✓, and differentiating gives M′(0) = r/p — r successes at 1/p trials each.
- (b) MGF of Y by the linear-change rule (a = 2p, b = 0): MY(t) = MX(2pt) = [pe2pt/(1 − qe2pt)]r.
- Divide top and bottom by p (both go to 0 — that's the problem to fix). The denominator 1 − (1−p)e2pt = (1 − e2pt) + pe2pt, so the bracket becomes e2pt / [ (1 − e2pt)/p + e2pt ].
- Take p → 0: e2pt → 1, and with h = 2pt, (1 − eh)/p = −2t · (eh − 1)/h → −2t. The bracket → 1/(1 − 2t), so MY(t) → (1 − 2t)−r, t < ½.
- Name it (recognition table): Gamma(α = r, β = 2), which is chi-squared with 2r degrees of freedom. Numerically, at r = 3, t = 0.2: MY = 3.860, 4.452, 4.611, 4.628 for p = 0.5, 0.1, 0.01, 0.001, closing on 0.6−3 = 4.630.
Second example, same move (binomial → Poisson). Bin(n, μ/n) has MGF (q + pet)n = (1 + p(et − 1))n = (1 + μ(et − 1)/n)n (using q = 1 − p). The second school limit, with a = μ(et − 1), sends it to eμ(et−1) — the Poisson(μ) MGF. That is why Module 2's "large n, small p" approximation works.
4 · Chebyshev's inequality L19 R1 §4.5
A guarantee that needs only μ and σ — no distribution. For any random variable and any k > 0:
In words: at most 1/k² of the probability lies k or more SDs from the mean; at least 1 − 1/k² lies strictly within. The price of knowing nothing about the shape is that the bound is loose — here it is against three real distributions (the handout says no proof, so none is given):
| k | Chebyshev bound 1/k² | Normal (exact) | Exponential (exact) | Uniform (exact) |
|---|---|---|---|---|
| 1.5 | 0.4444 | 0.1336 | 0.0821 | 0.1340 |
| 2 | 0.2500 | 0.0455 | 0.0498 | 0 |
| 3 | 0.1111 | 0.0027 | 0.0183 | 0 |
Each column is P(|X − μ| ≥ kσ). Exponential (any λ): only the right tail exists for k ≥ 1, giving e−(1+k). Uniform on [0, 1]: σ = 1/√12 ≈ 0.289, so it never reaches √3 ≈ 1.73 SDs from its mean.
Worked example 5Chebyshev, both directions
A plant's daily output has μ = 500 units and σ = 20; the distribution is unknown. (a) Lower bound for P(440 < X < 560). (b) An interval centred at μ that holds at least 96% of days. (c) Lower bound for P(450 < X < 560).
- (a) 440 and 560 are both 60 away from 500, and 60 = kσ gives k = 3. So P(|X − 500| < 60) ≥ 1 − 1/9 = 8/9 ≈ 0.8889. (If X happened to be normal the truth would be 0.9973 — the bound is a floor, not an estimate.)
- (b) Backwards: 1 − 1/k² = 0.96 ⟹ 1/k² = 0.04 ⟹ k = 5. Interval 500 ± 5(20) = (400, 600).
- (c) Not symmetric about 500 (50 below, 60 above) — Chebyshev only speaks about symmetric intervals. Use the largest symmetric interval inside it: (450, 550), k = 50/20 = 2.5. Since (450, 550) ⊂ (450, 560), P(450 < X < 560) ≥ P(|X − 500| < 50) ≥ 1 − 1/6.25 = 0.84.
5 · Transformation methods L20 R1 §6.7
Covered here: one random variable, Y = g(X), by the cdf method, the change-of-variable (pdf) formula for monotone g, and the MGF method. If your lecture also did two-variable transformations (Jacobians), that belongs with Module 4's joint densities — ask the instructor whether it is in mid-sem scope.
The question: you know the distribution of X; what is the distribution of Y = g(X)? The cdf method works every time — go through probabilities, never through densities directly:
- Write FY(y) = P(Y ≤ y) = P(g(X) ≤ y).
- Solve the inequality g(X) ≤ y for X — an interval (or intervals) of x-values. If g is decreasing, the inequality flips.
- Evaluate that probability with FX, then differentiate: fY = FY′. State the range of y.
Worked example 6cdf method, twice
(a) U ~ Uniform(0, 1), Y = −ln U. (b) Z ~ N(0, 1), Y = Z².
- (a) Range: U is in (0, 1), so ln U < 0 and Y > 0.
- (a) Solve: −ln U ≤ y ⟺ ln U ≥ −y (multiplying by −1 flips it) ⟺ U ≥ e−y (e(·) is increasing, no flip).
- (a) Evaluate: for a uniform, probability = length, so P(U ≥ e−y) = 1 − e−y. That is the Exp(1) cdf — Y ~ Exp(1), no differentiation needed. (This is how computers turn uniform random numbers into exponential ones.)
- (b) Solve: for y > 0, Z² ≤ y ⟺ −√y ≤ Z ≤ √y — two sides, because z² is not monotone (figure below). So FY(y) = Φ(√y) − Φ(−√y) = 2Φ(√y) − 1.
- (b) Differentiate (chain rule; Φ′ = φ, the N(0,1) density; (√y)′ = 1/(2√y)): fY(y) = 2φ(√y) · 12√y = e−y/2√(2πy), y > 0.
- (b) Name it: the gamma pdf with α = ½, β = 2 is y−1/2e−y/2/(Γ(½)·21/2), and Γ(½) = √π, so the denominators agree: Z² ~ χ²(1). The square of a standard normal is chi-squared with one degree of freedom.
The pdf formula (monotone g only)
When g is strictly increasing or strictly decreasing, the cdf method always produces the same shape of answer, so it can be packaged. Write x = h(y) for the inverse (solve y = g(x) for x). Then
Why the |h′|: a thin strip of x-values and its image strip of y-values hold the same probability, but if g stretches the strip, the same probability is spread over more width, so the density must drop by the stretch factor. The absolute value handles decreasing g (a density can't be negative).
Worked example 7The formula, and the MGF method
(a) X ~ Uniform(0, 1), Y = X²: find fY. (b) X ~ N(μ, σ²), Y = aX + b: find the distribution of Y.
- (a) On (0, 1), x² is increasing — formula allowed. Inverse h(y) = √y, h′(y) = 1/(2√y), and fX = 1 on (0, 1). So fY(y) = 1 · 1/(2√y) for 0 < y < 1 (the range: x ∈ (0, 1) maps to y ∈ (0, 1)).
- (a) Check: ∫01 12√y dy = [√y]01 = 1 ✓, and E(Y) = ∫01 y/(2√y) dy = ∫01 √y/2 dy = 1/3, matching E(X²) = ∫01 x² dx = 1/3 ✓.
- (b) MGF method — compute MY, then recognise it. Linear-change rule: MY(t) = ebtMX(at) = ebteμat + σ²a²t²/2 = e(aμ+b)t + (a²σ²)t²/2.
- (b) Recognise: normal shape, so Y ~ N(aμ + b, a²σ²). A linear function of a normal is normal — the fact behind every z-score (a = 1/σ, b = −μ/σ gives N(0, 1)).
1) Reading σ instead of σ² off a normal MGF: the t² coefficient is σ²/2 — double it for the variance, square-root for σ. In the 2025 Q1 shape, multiply the 0.5 through first. 2) Dropping the range of t (t < λ, t < 1/β, t < −ln q) — it's part of the answer. 3) Chebyshev as a value: it gives a bound ("at least 0.8889"), never "= 0.8889"; and for k ≤ 1 it says nothing (1/k² ≥ 1). 4) Asymmetric intervals in Chebyshev: shrink to the largest symmetric interval inside (WE5c). 5) Forgetting to flip an inequality when g is decreasing, or using the pdf formula on a non-monotone g like z² on the whole line — use the cdf method there. 6) No range for Y: every transformed pdf needs "for … < y < …".
Your minimal prerequisite kit for this part
- Exponent laws: eaeb = ea+b — every MGF derivation starts by merging exponents.
- Completing the square in one variable (shown before Worked example 2).
- Chain rule on (λ − t)−1, eg(t), Φ(√y). Drill: MATH U101 daily drill.
- ∫0∞ e−cx dx = 1/c, only for c > 0.
- Two limits: (eh − 1)/h → 1; (1 + a/n)n → ea.
- Inequalities: multiplying by a negative flips them; increasing functions (exp, ln, √ on positives) preserve them.
What to practise
Devore has no sections for these topics, so the practice sources are different from Modules 1–3:
| Skill | Where | How many |
|---|---|---|
| Derive a continuous MGF (exponential, gamma, uniform) and extract E, V | R1 (Milton & Arnold) exercises on MGFs in the continuous-distributions chapter; your tutorial sheets | 3–4 |
| Recognise a distribution from its MGF, then compute a probability | the mid-sem PYQ page (2025 Q1) and mid-sem drill | 4–6 |
| MGF limits (nb → gamma, binomial → Poisson) | 2025 Q7 on the PYQ page; re-derive WE4 without looking | 2 |
| Chebyshev: bound, backwards for k, asymmetric intervals | R1 §4.5 exercises; Devore §3.3's exercise set has Chebyshev problems (search the set for "Chebyshev") | 3–4 |
| cdf method and the monotone pdf formula | R1 §6.7 exercises | 4–5 |
Exercise numbering varies between printings — choose by section and skill. Where R1 isn't in hand, the lesson's check-yourself boxes and the drill are enough for the depth these topics have shown on past papers.