MATH U113 · First-time lesson · Module 3 · Part 2

MGFs, Chebyshev and transformations, taught from zero

The last three lectured topics of Module 3 — moment generating functions for continuous rvs (L12–16), Chebyshev's inequality (L19, R1 §4.5) and transformation methods (L20, R1 §6.7). None is in Devore. Budget about 2 hours; section 2 earns the biggest slice, because it's where past mid-sem marks were.

How to use this page

First time through → this page, in order, trying each green check before opening it. Revising → the notes page, which has the full worked examples and the MGF catalogue. Two facts to hold onto: this material is new to every person in the hall — it's not in Devore, not on any school syllabus, and JEE doesn't touch it; and it rests on things you already have — Module 2's MGF idea, Module 3's densities, the chain rule. What's new is three short ideas, one per section below.

1 · A barcode for curves L12–16

In Module 2's lesson the moment generating function was a barcode: one function, M(t) = E(etX), that carries every moment of X and identifies its distribution. Differentiate and set t = 0: out comes the mean; once more, E(X²).

For a continuous random variable, nothing about the idea changes. Only the arithmetic does, the same way it did all through Module 3: the sum over values becomes an integral over the density.

M(t) = E(etX) = ∫ etx f(x) dx

Take the friendliest density, the exponential, f(x) = λe−λx for x ≥ 0. Put it in, and the two exponentials merge into one, since etx · e−λx = e−(λ−t)x:

M(t) = λ ∫0∞ e−(λ−t)x dx

Here is the one genuinely new wrinkle. That integral runs to infinity, and it only has a finite value if the integrand dies away — if λ − t is positive. Try t = 2λ: the integrand becomes e+λx, which grows forever, and the area is infinite. So the barcode only exists for t < λ. That's no problem — we only ever use t near 0 — but the range is part of the answer, and you should write it down every time.

For t < λ, the integral of e−cx from 0 to ∞ is 1/c, so

M(t) = λλ − t, t < λ

Check the barcode's one universal feature: M(0) = λ/λ = 1 ✓ (it's the total probability). Differentiate: M′(t) = λ/(λ − t)², so M′(0) = 1/λ — the exponential's mean, as promised. The notes page finishes the variance and does the gamma the same way; its table lists all five MGFs you should know by heart for a closed-book paper.

Check yourself: X is uniform on [0, 1]. Find M(t) for t ≠ 0.

f = 1 on [0, 1], so M(t) = ∫01 etx dx = [etx/t]01 = (et − 1)/t. At t = 0 the formula reads 0/0, which is why the table says "and 1 at t = 0" — every MGF equals 1 there. No range restriction: the interval is finite, so the integral always exists.

Check yourself: X has MGF (1 − 3t)⁻⁴. What distribution, and what is E(X)?

The shape (1 − βt)−α is a gamma's, with β = 3, α = 4. Mean αβ = 12. By differentiating instead: M′(t) = −4(1 − 3t)−5 · (−3) = 12(1 − 3t)−5, and at 0 that is 12 ✓.

2 · The normal's barcode, and reading it backwards

The normal's MGF is the one exams love, and its derivation hinges on one piece of algebra worth doing slowly. For the standard normal Z, the integral is

MZ(t) = ∫ 1√(2π) etz − z²/2 dz

The exponent, tz − z²/2, is almost the exponent of a bell curve, −(z − something)²/2. Completing the square makes "almost" exact. Start from the square we want and expand it: (z − t)² = z² − 2tz + t². Halve it and put a minus sign in front: −(z − t)²/2 = −z²/2 + tz − t²/2. Compare with our exponent: it's the same except for the extra −t²/2. So add that back:

tz − z²2 = −(z − t)²2 + t²2

Now et²/2 doesn't involve z, so it walks out of the integral, and what stays inside is the density of a normal centred at t — the standard bell, slid sideways. A density's total area is 1. So MZ(t) = et²/2. Stretching by σ and shifting by μ (X = μ + σZ) turns it into

MX(t) = eμt + σ²t²/2

(the notes show the stretch-and-shift rule in one line).

Now the move that actually earns marks: reading it backwards. Because an MGF identifies its distribution uniquely, an exam can hand you an MGF instead of naming the distribution. Your job is to spot the shape and read the parameters off it. For a normal: the coefficient of t is the mean; the coefficient of t² is half the variance.

The 2025 mid-sem opened exactly like this, for 9 marks: exam-completion time has MGF e0.5(σ²t² + 160t). The 0.5 in front is disguising it. Multiply through: 80t + σ²t²/2. Normal, with mean 80. The rest of the question (finding σ from a given tail probability, then two table lookups) is pure §4.3 — the notes' Worked example 3 does it in full, with a figure.

Check yourself: M(t) = e^(3t + 8t²). What is the distribution? What is P(X > 7)?

Normal shape. Coefficient of t: μ = 3. Coefficient of t² is σ²/2 = 8, so σ² = 16, σ = 4. Then P(X > 7) = 1 − Φ((7 − 3)/4) = 1 − Φ(1) = 1 − 0.8413 = 0.1587. The classic slip is taking σ = 8 or σ² = 8.

Check yourself: M(t) = 1/(1 − 2t)⁵. Name the distribution two ways.

(1 − 2t)−5 is gamma with α = 5, β = 2. A gamma with β = 2 is a chi-squared with ν = 2α = 10 degrees of freedom. Mean 10, variance 20.

3 · When barcodes converge

Uniqueness has a second, more surprising use. Suppose you have a whole family of random variables — one for each value of some parameter — and you want to know what the distribution looks like as the parameter goes to an extreme. Computing the pmfs and taking a limit is usually hopeless. But if the MGFs settle down to the MGF of a known distribution, the distributions settle down to it too. So you take the limit of one formula instead of a whole distribution.

The 2025 mid-sem's final question (12 marks) was exactly this. X counts the trials needed for the r-th success; its MGF is [pet/(1 − qet)]r. When p is tiny, X is huge (you wait ages for each success), so the question rescales it: Y = 2pX. Watch what happens to Y's MGF as you drag p toward 0:

dashed: limit (1 − 2t)^(−r) solid: MGF of Y = 2pX

The solid curve climbs onto the dashed one. The dashed curve is (1 − 2t)−r — from section 2's second check, that's a gamma with α = r, β = 2, which is chi-squared with 2r degrees of freedom. So for small p, 2pX is approximately χ²(2r).

On paper, the limit needs one fact from school calculus:

eh − 1h → 1 as h → 0

It's the slope of ex at x = 0 written as a limit (rise eh − e0 over run h), and that slope is e0 = 1. The trouble in Y's MGF is that numerator and denominator both go to 0 as p → 0. Dividing both by p turns the denominator into a piece of exactly that shape, and it tends to 1 − 2t. The notes' Worked example 4 does it line by line, including part (a)'s MGF derivation.

How deep to go

The handout doesn't name "limits via MGFs", but the 2025 paper spent 12 marks on it. Learn this one example and the binomial → Poisson version in the notes; confirm with your instructor before going further.

Check yourself: at p = 0 exactly, why can't you just substitute into Y's MGF?

Substituting p = 0 into pe2pt/(1 − (1 − p)e2pt) gives 0/(1 − 1) = 0/0 — no information. That is exactly why you divide top and bottom by p first, then take the limit. (And physically, at p = 0 there are no successes at all, so X doesn't exist; only the limit makes sense.)

4 · Chebyshev: a guarantee from two numbers L19

Everything so far assumed you know the distribution. Suppose you don't. A supplier tells you only that their bolts have mean length 50 mm and SD 2 mm. Can you say anything about how many bolts lie far from 50?

Surprisingly, yes. Chebyshev's inequality says that for any distribution at all, and any k > 0:

P(|X − μ| ≥ kσ) ≤ 1k²

Read it as: at most 1/k² of the probability can be k or more standard deviations from the mean. At most a quarter lies 2 SDs out; at most a ninth lies 3 SDs out. Flip it and you get a floor: at least 1 − 1/k² lies strictly within k SDs. For the bolts, at least 1 − 1/4 = 75% are within 46–54 mm, whatever the distribution is. (The handout asks for the statement without proof.)

The price of knowing so little is that the guarantee is usually loose. Slide k and switch shapes to see how far the true tail sits below the bound:

distance from μ, in standard deviations

Three things to notice. The shaded area is always below the bound — that's the theorem. For k ≤ 1 the bound is 1 or more, so it says nothing at all. And the exponential, which is lopsided, has no left tail beyond 1 SD (it can't go below 0), yet Chebyshev still holds for it — it holds for everything.

Exam questions run it forwards ("at least what fraction lies between 440 and 560?") or backwards ("how wide must the interval be to hold 96%?" — solve 1 − 1/k² = 0.96, giving k = 5). The notes' Worked example 5 does both, plus the trap of an interval that isn't centred on μ.

Check yourself: μ = 50, σ = 2. At least what fraction of bolts lies strictly between 44 and 56 mm?

44 and 56 are 6 away from 50, and 6 = 3 × 2, so k = 3: at least 1 − 1/9 = 8/9 ≈ 0.8889. Write "at least" — Chebyshev never gives an exact value.

5 · Transformations: new distributions from old L20

Last idea. You know how X is distributed; someone computes a new quantity from it — Y = X², or Y = −ln X, or a temperature converted from °C to °F. How is Y distributed?

The tempting move — plug g into the density — is wrong, because a density isn't a probability. The safe route goes through something that is a probability, the cdf. That's the whole cdf method:

  1. Write FY(y) = P(Y ≤ y) and replace Y by its formula in X.
  2. Solve that inequality for X. (Careful: multiplying by a negative number flips it.)
  3. You now have a probability about X — compute it. Differentiate if you want the density.

Try it on a famous one. U is uniform on (0, 1) — a calculator's random-number button — and Y = −ln U. Since U < 1, ln U is negative, so Y is positive. Then Y ≤ y means −ln U ≤ y, which means ln U ≥ −y (the minus flipped it), which means U ≥ e−y. For a uniform on (0, 1), the probability of landing in an interval is its length, so P(U ≥ e−y) = 1 − e−y. That's the exponential cdf with λ = 1. The random-number button, passed through −ln, produces exponential waiting times — this is how simulations generate them.

One shape needs extra care: a g that goes down and then up, like z². Then "Y ≤ y" catches values of Z on both sides of 0:

−√y √y 0 y = z² level y z
Every z in the blue band has z² below the dashed level, so P(Z² ≤ y) = P(−√y ≤ Z ≤ √y) — both arms of the parabola count.

For Z ~ N(0, 1) that gives FY(y) = 2Φ(√y) − 1, and differentiating turns out to give the chi-squared density with 1 degree of freedom: the square of a standard normal is χ²(1). That's the chi-squared's origin story, and why it keeps appearing in the statistics half of the course. (Full derivation: notes, Worked example 6.)

When g only ever increases (or only decreases), the cdf method always produces the same shape of answer, which the notes package as a formula: fY(y) = fX(x) · |dx/dy|, with x written in terms of y. The |dx/dy| factor corrects for stretching: if g spreads an interval out, the same probability covers more width, so the density there must be lower.

Check yourself: X ~ Exp(λ) and Y = 2X. Use the cdf method to find Y's distribution.

FY(y) = P(2X ≤ y) = P(X ≤ y/2) = 1 − e−λy/2 for y ≥ 0. That's the exponential cdf with rate λ/2: doubling a waiting time halves the rate and doubles the mean (2/λ), as it should.

Check yourself: X ~ Uniform(0, 1) and Y = X². What is P(Y ≤ 0.25)?

X² ≤ 0.25 ⟺ X ≤ 0.5 (on (0, 1), X is positive, so only one side counts). P(X ≤ 0.5) = 0.5. So half of Y's probability is crammed into [0, 0.25] — squaring piles small values up near 0.

6 · You're ready — what to do next

The part in one sentence: an MGF is a barcode you can derive (an integral), read backwards (uniqueness) and take limits of; Chebyshev bounds tails using only μ and σ; and the cdf method turns "distribution of g(X)" into a probability about X. Next:

Still stuck on something? Ask an AI well

Pin it to the syllabus and make it interactive:

"I'm studying moment generating functions of continuous random variables (Milton & Arnold, as used in my Probability & Statistics course). Give me an MGF, one at a time — normal, gamma, exponential or chi-squared, sometimes disguised with a constant multiplied through — and ask me to name the distribution and its parameters. Check my answer before giving the next one."

"Quiz me on Chebyshev's inequality P(|X − μ| ≥ kσ) ≤ 1/k²: alternate between 'find the lower bound for this interval' (sometimes not centred on μ) and 'find the interval that guarantees this fraction'. One question at a time; check my k before continuing."

"Walk me through the cdf method for Y = g(X) on one example where g is decreasing, asking me to do each step (write F_Y, solve the inequality, evaluate, differentiate) before you show it."

One caution: AI answers can contain confident arithmetic errors — recompute any final number yourself.