MATH U113 · First-time lesson · Module 2

Discrete random variables, taught from zero

Devore (9th ed., Metric) §3.1–3.6 — random variables, pmf and cdf, expected value and variance, then the four named models: binomial, hypergeometric, negative binomial, Poisson. Budget 2–3 hours across sittings; sections 4–6 are the exam heart.

How to use this page

First time through → this page, in order, attempting each green check before opening it. Revising → the notes page. Status check: this chapter is the start of the distribution framework — the way of thinking the whole course is really about — and it is new to everyone. JEE classmates met the binomial formula; the framework around it (pmf, cdf, expectation as long-run average, model choice) is fresh territory for the entire hall. Two facts from the handout: this is the course's biggest lecture block (L3–11, nine lectures) and prime midsem material; and the syllabus adds moment generating functions from the reference book (Milton & Arnold R1 §3.4) — not in Devore, so section 3½ below teaches it and the notes page carries the derivations; treat lecture notes as the primary source for exam style. Everything here also leans on Module 1 — especially counting and independence — so revise that first if it's shaky.

1 · What a random variable is §3.1

Module 1 worked with events — verbal things like "at least one head". Doing arithmetic with words is clumsy, and engineering questions are numerical anyway: how many defectives, how long until failure, how much current. So we attach a number to every outcome.

A random variable (rv) is a rule that assigns one number to each outcome of the sample space. Toss a coin twice and let X = number of heads:

HH → 2 HT → 1 TH → 1 TT → 0

Before the experiment, X's value is uncertain — that's the "random". After, it's a plain number. Notation: capital X for the variable, lowercase x for a value it might take; "X = 1" is an event (here {HT, TH}), so it has a probability. That one move — events become equations — is the whole point of the chapter.

Two species: a discrete rv takes isolated values you can list (0, 1, 2, …) — counts, mostly. A continuous rv fills whole intervals — times, lengths, currents. This chapter is the discrete story; Chapter 4 retells it for continuous (and the parallel is deliberate — learn the pattern here and Chapter 4 is half-familiar on arrival). The simplest rv of all, the Bernoulli rv, takes only values 0 and 1 ("failure"/"success") — it looks trivial, but section 4 builds its biggest model out of stacks of these atoms.

2 · The distribution: pmf and cdf §3.2

To know a discrete rv completely, you need exactly two columns: what values can it take, and with what probabilities. That table is the probability mass function (pmf), p(x) = P(X = x). For the two-toss X (four equally likely outcomes):

x012
p(x)0.250.500.25

Every legitimate pmf obeys two laws, straight from Module 1's axioms: each p(x) ≥ 0, and they sum to 1 (some value must occur). That second law is an exam workhorse: "find k so this is a valid pmf" means "make the column sum to 1". Drawn as a bar chart (a probability histogram), the pmf is literally the idealised shape that Module 1's data histograms approach as samples grow — the two halves of the course shaking hands.

The second description is the cumulative distribution function (cdf): the running total,

F(x) = P(X ≤ x) = Σy ≤ x p(y)

For the table above: F(0) = 0.25, F(1) = 0.75, F(2) = 1 — and between listed values it holds flat, a staircase that starts at 0 and climbs to 1. Why bother with a running total? Because textbook tables and software give you cdfs, and every interval probability comes from two lookups:

P(a ≤ X ≤ b) = F(b) − F(a − 1)  (integer-valued rvs)

The a − 1 is not a typo: F(b) already includes a itself, so subtract only what's strictly below a. Off-by-one slips here are the single biggest mark-loser of the chapter — always translate the words into an inequality first ("at least 8" → X ≥ 8 → 1 − F(7)), then reach for the table.

Check yourself: X takes values 1, 2, 3, 4 with p(1) = 0.2, p(2) = 0.3, p(3) = k, p(4) = 0.1. Find k, then F(2), then P(X ≥ 3).

Sum to 1: 0.2 + 0.3 + k + 0.1 = 1 → k = 0.4. F(2) = 0.2 + 0.3 = 0.5. P(X ≥ 3) = 1 − F(2) = 0.5 (or directly 0.4 + 0.1).

Building a distribution from scratch: the maximum of two dice

So far the pmf table was handed to you. Tutorial questions hand you an experiment instead and ask you to build the table yourself. Here is the standard one, worked from absolute zero — this is the "T3 Q3" pattern, and once you've seen it built once, the whole family (max, min, sum, difference) opens up.

The setup. Toss two fair dice independently. Define M = the larger of the two numbers (if they tie, that shared number). So if the dice show 1 and 5, M = 5; if both show 3, M = 3. Note what kind of thing M is: not a die, but a rule that attaches one number to each outcome of the experiment — a random variable, exactly as in section 1.

What the question is actually asking. Part (a), "the pmf of M": pmf is short for probability mass function, and it is nothing more than the complete table for M — every value it can take (here 1…6), and the probability of each. Deliver that table and part (a) is answered in full. Part (b), "the cdf": the cumulative distribution function, F(m) = P(M ≤ m) — the running total of the same table, "the probability the max is at most m". (Section 2 above defines both properly with a simpler example — if the words are brand new, read that first and come back.) So the entire problem is: build one table, then add it up cumulatively. Here is how.

Step 1: write down the sample space — it's a grid. The experiment's outcome is a pair: what die A shows and what die B shows. Six options each, so 6 × 6 = 36 pairs, and because the dice are fair and independent, all 36 pairs are equally likely — each has probability 1/36. Every two-dice problem starts with this grid; drawing it (even quickly, even just 6×6 dots) is not a childish step, it's the professional step.

Step 2: mark where your variable takes each value. One notation reminder before the counting (same convention as section 1's X vs x): capital M is the random quantity itself — unknown until the dice land; lowercase m is a blank standing for one particular value it might take, one of 1…6. So "p(m) = P(M = m)" is a question you ask six times, once per value of the blank. Ask it first with the blank set to 4: which cells of the grid have M = 4? You need at least one die showing 4 and neither die above 4. Slide the control and watch the shape:

How to read the grid: each cell is one outcome — its row number is what die A shows, its column number what die B shows — and the number written inside the cell is the value of M for that outcome, i.e. the larger of its row and column numbers. Check one by hand: the cell in row 2, column 5 is the outcome "A shows 2, B shows 5", and it's labelled 5 because max(2, 5) = 5. The diagonal cells are the ties — (3,3) is labelled 3. Fill in three or four cells yourself before reading on; the labels only mean something once you've computed a few.

cells where M = m — the L-shaped band · cells where M < m — the square it wraps around

The cells with M = 4 form an L-shaped band: the row where A = 4 (with B = 1…4) and the column where B = 4 (with A = 1…4). That looks like 4 + 4 = 8 cells — but the corner cell (4, 4) sits in both the row and the column, and it's one cell, counted once. So: 4 + 4 − 1 = 7 cells. In general the band for M = m has m + m − 1 = 2m − 1 cells.

Step 3: repeat for every value, then divide by 36. The table isn't produced by a formula — it's produced by asking the step-2 question six times, once per value of m. Do the first two by hand: M = 1 needs both dice ≤ 1, so only the cell (1,1) — 1 cell. M = 2 is the row (2,1),(2,2) plus the column (1,2),(2,2), corner counted once — 3 cells. Slide the widget through m = 1…6 and read off the rest: each band adds two cells to the previous one (its row-arm and column-arm each grow by one), which is why the counts climb 1, 3, 5, 7, 9, 11 — and where the summary 2m − 1 comes from. Each count over 36 is one row of the pmf:

m123456
cells1357911
p(m)1/363/365/367/369/3611/36
That table IS the answer to part (a)

Nothing more is coming — the pmf of M is exactly this table: the values 1…6 with their probabilities 1/36, 3/36, 5/36, 7/36, 9/36, 11/36. On an exam, writing this table (or the one-line summary p(m) = (2m−1)/36 for m = 1,…,6, which step 5's shortcut derives) is full marks for "what is the pmf of M?".

Step 4: audit before moving on. The counts 1 + 3 + 5 + 7 + 9 + 11 = 36 ✓ — every cell of the grid claimed exactly once, so the probabilities sum to 1, as section 2's law demands. And the shape makes sense: a maximum gets dragged upward, so big values of M should be more likely than small ones — and 11/36 > 1/36 indeed. If your pmf doesn't sum to 1, you double-counted a corner or missed a band; the grid shows you where.

Step 5: the cdf is the running total — F(m) = P(M ≤ m), adding the pmf left to right:

m123456
F(m)1/364/369/3616/3625/3636/36

Stare at that row: 1, 4, 9, 16, 25, 36 — perfect squares. F(m) = m2/36. The grid explains why: "M ≤ m" says the larger die is at most m — which is just a wordier way of saying both dice are at most m. And "both ≤ m" is the m × m square in the grid's corner: m2 cells (the green square plus the blue band in the widget). No adding needed.

Because it's a discrete variable, the full cdf is a staircase: 0 for m < 1, flat at m2/36 between integers, jumping by p(m) at each integer, reaching 1 at m = 6 and staying there. When an exam asks to "determine the cdf", write the formula and say it holds as a step function (or sketch the staircase) — that's the complete answer.

And that's part (b) answered

The cdf of M is: F(m) = 0 for m < 1; F(m) = ⌊m⌋2/36 for 1 ≤ m < 6 (flat between integers — use the whole-number part); and F(m) = 1 for m ≥ 6. Writing the values 1/36, 4/36, 9/36, 16/36, 25/36, 1 at m = 1…6 plus the staircase sentence is the complete exam answer to "determine the cdf of M".

The professional shortcut (what the tutorial solution does)

Experienced solvers run this backwards: get F(m) = m2/36 first from the "both dice ≤ m" square — one line, no enumeration — then recover the pmf as consecutive differences, p(m) = F(m) − F(m−1) = (m2 − (m−1)2)/36 = (2m−1)/36 — the band is the square minus the smaller square. Same answer as the grid count. Both routes are full marks; the grid is the one you can always fall back on, the cdf-first route is the one that scales (max of three dice: F(m) = m3/216, no grid drawable).

Check yourself: L = the smaller of the two dice. Find P(L = 4) with the grid, then find F(m) — which corner does the square sit in now?

Cells with L = 4: at least one die shows 4, neither below 4 — the L-band anchored at the far corner: (4,4), (4,5), (4,6), (5,4), (6,4) → 5 cells, P = 5/36. General count 2(7−m) − 1 = 13 − 2m. For the cdf, the easy event is now "L ≥ m" = both dice ≥ m = the (7−m)2 square in the top corner, so F(m) = 1 − (6−m)2/36. Check: F(4) = 1 − 4/36 = 32/36, and 32 − F(3) = 32 − 27 = 5 ✓. Minimum problems run on the survival side — that's the whole difference.

Classic trap

Double-counting the corner: "M = 4 is a row of 4 plus a column of 4, so 8 cells." The cell (4,4) is one outcome, not two. Any time a count comes from two overlapping strips, subtract the overlap once — this is Module 1's inclusion–exclusion, back again. The sum-to-1 audit in step 4 is what catches it.

3 · Expected value and variance §3.3

Play a game a million times; what's your average result per play? Each value x shows up in a fraction p(x) of the plays (that's what probability means — Module 1's relative-frequency idea), so the long-run average is the probability-weighted sum:

E(X) = μ = Σ x · p(x)

For the two-toss rv: E(X) = 0(0.25) + 1(0.50) + 2(0.25) = 1 — one head expected in two tosses; sanity itself. It's Module 1's sample mean with relative frequencies hardened into probabilities, and the same balance-point picture applies to the probability histogram. Note E(X) needn't be a possible value: a family with 1.8 children on average has never met a 0.8 child. It's a long-run average, not a prediction for one trial.

Need the expected value of a function of X — a cost, say? Weight the function's values by the same pmf: E[h(X)] = Σ h(x) p(x). For the special case of a straight line, this collapses to the linearity rule you'll use constantly: E(aX + b) = a E(X) + b.

Spread transfers from the descriptive-statistics lesson the same way. The variance is the expected squared distance from the mean, with the same shortcut as before:

V(X) = σ2 = Σ (x − μ)2 p(x) = E(X2) − [E(X)]2 · σ = √V(X)

Two-toss rv again: E(X2) = 0(0.25) + 1(0.50) + 4(0.25) = 1.5, so V(X) = 1.5 − 1² = 0.5 and σ ≈ 0.707. (Note in passing: E(X2) = 1.5 ≠ 1 = [E(X)]² — squaring and averaging do not commute; the gap between them is the variance.) And the rescaling rules: V(aX + b) = a²V(X) — shifting by b moves the whole distribution rigidly (no spread change, so b vanishes), and scaling by a stretches distances, which squared distances feel as a².

Check yourself: a ₹10 lottery ticket wins ₹100 with probability 0.05 (else nothing). Let X = net gain. E(X)?

X = +90 with probability 0.05, and −10 with probability 0.95. E(X) = 90(0.05) + (−10)(0.95) = 4.5 − 9.5 = −₹5. Long-run: you lose ₹5 per ticket on average — the mathematical definition of "the house always wins".

3½ · The moment generating function: a barcode for distributions R1 §3.4

First, why anyone would invent this. You've just computed E(X) and E(X²) as separate sums — and for the named models coming in sections 4–6, those sums get genuinely ugly (try summing Σ x² e−μμx/x! directly). The moment generating function is a machine that does all such sums at once: build one function, and every E(Xk) falls out by differentiation.

The machine: M(t) = E(etX) — an expected value like any other, just of the odd-looking quantity etX, where t is a helper variable that means nothing by itself. Compute it once and it stores the entire distribution (hence "barcode": two rvs with the same MGF have the same distribution). To read the barcode, differentiate with respect to t, then set t = 0: the first derivative gives E(X), the second gives E(X²).

See it on the smallest possible case — a Bernoulli trial (X = 1 with probability p, else 0): M(t) = et·0q + et·1p = q + pet. Differentiate: M′(t) = pet, so M′(0) = p — which is indeed E(X). Differentiate again: M″(0) = p = E(X²), so σ² = p − p² = pq. The calculus load is exactly what MATH U101 is drilling right now: derivatives of eat, chain rule, product rule.

Check yourself 3½: a fair die shows 1 with probability 1/6, else 0 (call this X). Write M(t) and use it to get E(X).

M(t) = 5/6 + (1/6)et. Then M′(t) = (1/6)et and M′(0) = 1/6 — matching the direct computation E(X) = 0 · 5/6 + 1 · 1/6. The machine agrees with the sum; that's the whole contract.

The payoff scales: the notes page derives the binomial's MGF (one line via the binomial theorem) and the Poisson's, then extracts np and μ in two lines each — sums that take a page head-on. Exam questions are typically "derive the MGF of …" or "here is an MGF, identify the distribution / find mean and variance". The one habit that prevents every MGF disaster: differentiate first, set t = 0 second.

4 · The binomial distribution §3.4

Now the chapter's main event. An enormous number of situations share one skeleton, the binomial experiment:

Ten components tested (each passes with p = 0.9), twenty multiple-choice guesses (p = 0.25), five phone calls that connect or don't — same skeleton. Let X = number of successes. What's P(X = x)? Build it in two Module 1 moves:

  1. One specific sequence with x successes — say SS…SFF…F — has probability px(1−p)n−x, by independence (multiply along the sequence). Every other sequence with x successes has the same probability — multiplication doesn't care about order.
  2. How many such sequences? Choose which x of the n slots hold the successes: (n choose x) — §2.3 doing its job.
b(x; n, p) = (nx) px (1 − p)n−x · E(X) = np · V(X) = np(1−p)

(E(X) = np is exactly what instinct says: 10 trials at 90% → expect 9.) Play with the shape — this is the distribution you'll compute with most this semester:

μ x = number of successes

Three things to notice while sliding: at p = 0.5 the histogram is symmetric; away from 0.5 it skews (toward the rarer outcome's side); and as n grows, the shape smooths into a bell centred at np — that bell is Chapter 4's normal distribution announcing itself early.

Bookkeeping notes for exams: "at least/at most" questions use the cdf B(x; n, p) tabulated in the book's Appendix — with the off-by-one care from section 2. And the model-check matters more than the formula: if trials aren't independent or p drifts, it's not binomial, whatever it looks like — which is the cue for the next section.

Check yourself: fair coin, 4 tosses, P(exactly 2 heads)? (Set up b(x; n, p) first.)

b(2; 4, 0.5) = (4 choose 2)(0.5)²(0.5)² = 6/16 = 0.375. Not 0.5! Even the most likely count happens well under half the time — the rest spreads over 0, 1, 3, 4 heads.

One video, if you want one

Binomial distributions (3Blue1Brown, ~12 min) rebuilds the pmf exactly as above, animated. Watch once, after this section. At the end it pitches "probabilities of probabilities" (the beta distribution, part 2) — off-syllabus; stop there.

5 · The binomial's two cousins §3.5

Two model variations cover the situations the binomial checklist rejects. (Scope note: the handout confirms §3.5 is fully in-syllabus — the negative binomial, geometric and hypergeometric all get lecture time.)

Hypergeometric: sampling without replacement

Drawing n items from a small population of N items, M of them "successes", without replacement: each draw changes the pool, so trials aren't independent and p isn't constant — not binomial. But you've already computed with this model! Module 1's Worked Example 1 (4 boards from 12, 3 defective) was the hypergeometric pmf, before it had a name:

h(x; n, M, N) = (M choose x) (N−M choose n−x)(N choose n) · E(X) = n · MN

The mean is again "instinct's answer" (sample 4 from a pool that's ¼ defective → expect 1). When the population is huge relative to the sample (n/N ≤ 0.05 as a working rule), removing a few items barely changes the pool, and the binomial is an excellent stand-in — that's why "2 phones from a warehouse of 40,000" gets treated as binomial without apology.

Negative binomial: trials until enough successes

Flip the design: don't fix the number of trials — keep going until the r-th success, and let X = number of failures along the way (Devore's convention; some books count total trials — state whose you're using). The r = 1 case is the geometric distribution: failures before the first success. Its headline fact is the intuitive one: expected failures before the first success = (1−p)/p — for a 10%-chance event, expect about 9 misses before the hit, i.e. success around trial 10 on average.

6 · The Poisson distribution: counting rare events §3.6

How many calls hit a helpdesk in an hour? Flaws per 100 m of cable? Particles decaying per second? There's no fixed n of "trials" here — just a rate and a window. The model for counts of rare, independent events is the Poisson distribution:

p(x; μ) = e−μ μxx!, x = 0, 1, 2, … · E(X) = V(X) = μ

Where does it come from? It's the binomial pushed to its limit: slice the hour into thousands of instants, each a "trial" with a tiny success chance. As n → ∞ and p → 0 with np → μ held fixed, b(x; n, p) → p(x; μ). Practically that gives you two uses: a model in its own right for event counts, and an approximation to the binomial when n is large and p tiny (rule of thumb: n ≥ 50, np ≤ 5).

The one operational rule that decides most exam questions: μ must be the rate scaled to the window asked about. Calls arrive at α = 2 per minute and the question asks about 3 minutes? Then μ = αt = 6, not 2. (This "events in time at rate α" setting is the Poisson process; its inter-arrival times become Chapter 4's exponential distribution — another handshake across chapters.) A remarkable signature worth remembering: mean and variance are equal — count data whose sample mean ≈ sample variance whispers "Poisson".

Check yourself: a proofreader finds typos at rate 0.5 per page. P(a given page has no typos)?

μ = 0.5 for one page. P(X = 0) = e−0.5(0.5)⁰/0! = e−0.5 ≈ 0.607. About 6 pages in 10 are clean — even at half a typo per page.

7 · You're ready — what to do next

The chapter in one sentence: a discrete rv is fully described by its pmf (or cdf); expectation and variance compress it to centre and spread; and four named pmfs — binomial, hypergeometric, negative binomial, Poisson — cover most counting situations, chosen by checking the experiment's structure, not by formula-guessing. Next:

Still stuck on something? Ask an AI well

Pin it to the syllabus and make it interactive:

"I'm studying Devore 9th ed. §3.4–3.6. Give me 6 short scenarios, one at a time; I say which model applies (binomial / hypergeometric / negative binomial / Poisson / none) and why, citing the model's conditions — correct me before the next one. Don't make me compute yet."

"Walk me through computing E(X) and V(X) from a pmf table (Devore §3.3 level) using both the definition and the E(X²) − μ² shortcut, then give me one table to do alone and check both my answers."

One caution: AI answers can contain confident arithmetic errors — recompute any final number yourself.