MATH U113 · First-time lesson · Module 1
Probability, taught from zero
Devore (9th ed., Metric) §2.1–2.5 — the language of chance, the three axioms, counting, conditional probability, and independence. Budget 2–3 hours in two or three sittings; section 3 (counting) deserves its own sitting.
First time through → this page, in order, attempting every green check box before opening it. Revising → the notes page. One honest note up front: section 3 (counting) is the single place in this course where JEE-drilled classmates genuinely have a head start — it's a JEE staple and 11th-class material. That's why this lesson teaches it from absolute zero and why it gets the most check boxes. It is one section wide; work it slowly once, and the gap is closed. And a fact from the handout worth knowing: §2.1 and §2.3 are official self-study — no lecture covers them, for anyone. The whole hall teaches themselves counting; you're doing it with a lesson built for exactly that, while classmates lean on two-year-old JEE memory. The lectured sections (§2.2, 2.4, 2.5) take just two lectures, L1–2 — read sections 2, 4 and 5 here before those lectures, not after.
1 · The language: experiments, sample spaces, events §2.1
Probability needs a precise vocabulary before it can say anything. Three words do all the work:
An experiment is any activity with an uncertain outcome — toss a coin twice, measure a resistor, count today's server crashes. The sample space 𝒮 is the list of every possible outcome. Toss a coin twice and there are exactly four:
(Note HT ≠ TH — the order of listing is part of the outcome. Writing the sample space carelessly is where most beginner errors are born.)
An event is any collection of outcomes — any subset of 𝒮. "At least one head" is the event A = {HH, HT, TH}. An event occurs if the actual outcome is one of its members. Single outcomes like {HH} are simple events; everything else is compound.
Because events are sets, the three English connectives become three set operations — and this dictionary is used on every page of the course:
| English | Set notation | Meaning |
|---|---|---|
| A or B (or both) | A ∪ B (union) | outcomes in either |
| A and B | A ∩ B (intersection) | outcomes in both |
| not A | A′ (complement) | outcomes not in A |
Two events with no outcomes in common — A ∩ B = ∅ — are mutually exclusive (or disjoint): they cannot both happen. "First toss is H" and "first toss is T" are mutually exclusive; "first toss is H" and "at least one T" are not (HT is in both). Draw a Venn diagram whenever the algebra gets tangled — two overlapping circles resolve most §2.1 exercises on sight.
Check yourself: a die is rolled once. Write the sample space, the event E = "even", the event L = "at most 2", then E ∩ L, E ∪ L, and E′.
𝒮 = {1, 2, 3, 4, 5, 6}; E = {2, 4, 6}; L = {1, 2}. Then E ∩ L = {2}, E ∪ L = {1, 2, 4, 6}, E′ = {1, 3, 5}. Are E and L mutually exclusive? No — they share the outcome 2.
2 · The rules: three axioms and everything they buy §2.2
What is the probability of an event? The most useful answer is operational: the long-run relative frequency. Toss a fair coin 10 times and you might see 7 heads; toss it 10,000 times and the fraction of heads will be pinned near 0.5. Probability is the number the fraction settles on. Watch it happen — each run below is 500 fresh simulated tosses:
Notice the shape every run shares: wild swings early (after 3 tosses the fraction can only be 0, ⅓, ⅔ or 1), then an ever-calmer settling toward 0.5 — without the tosses themselves ever becoming predictable. That settling is what "P(heads) = 0.5" means. It's also why casinos and insurers, who cannot predict any single event, profit with near-certainty over thousands.
Rather than argue philosophy, the textbook distils everything a probability must satisfy into three axioms: for any event A, P(A) ≥ 0; the sure event gets P(𝒮) = 1; and for mutually exclusive events, probabilities add: P(A₁ ∪ A₂ ∪ ⋯) = P(A₁) + P(A₂) + ⋯. Everything else in Chapters 2–4 is a consequence. The three you'll use daily:
The last one — the addition rule — deserves one honest look at why: adding P(A) and P(B) counts the overlap twice (it sits inside both circles of the Venn diagram), so subtract it once. When A and B are mutually exclusive the overlap is empty and the rule collapses to plain addition — axiom 3 again.
And the complement rule is secretly the chapter's best labour-saving device: "the probability of at least one…" is almost always computed as 1 − P(none), because "none" is one clean event while "at least one" is a pile of cases.
One more consequence: when a sample space has N outcomes that are all equally likely (fair coins, fair dice, well-shuffled cards, items drawn "at random"), each outcome has probability 1/N, so for any event,
This is school's "favourable over total" — now with its fine print visible: only for equally likely outcomes. (A biased coin's sample space is still {H, T}; the formula would blindly say ½.) Its real consequence is strategic: probability questions become counting questions — how many outcomes are there, and how many are favourable? Which is why the next section exists.
Check yourself: 60% of hostel rooms have a fan issue (F), 30% a plumbing issue (M), 10% both. What fraction has at least one issue? Neither?
P(F ∪ M) = 0.6 + 0.3 − 0.1 = 0.8 — subtracting the double-counted overlap. Neither: complement rule, 1 − 0.8 = 0.2.
3 · Counting, from absolute zero §2.3
Everything here grows from one seed, the product rule: if a task happens in stages, and stage 1 can be done in n₁ ways and stage 2 in n₂ ways regardless of how stage 1 went, the pair can be done in n₁n₂ ways. Four shirts and three trousers make 4 × 3 = 12 outfits: for each shirt, all three trousers are available. Draw it as a tree — 4 branches, each sprouting 3 — and the rule is just "count the leaves". It extends to any number of stages by the same logic.
Two big counting tools are the product rule applied to one special task: selecting k objects from n distinct objects. Everything hinges on one question — does the order of selection matter?
Order matters: permutations
Ten sprinters; how many ways to fill gold–silver–bronze? Stage 1 (gold): 10 choices. Stage 2 (silver): 9 remain. Stage 3 (bronze): 8. Product rule:
In general, k ordered selections from n: start at n and count down k factors. The compact notation uses the factorial, n! = n(n−1)(n−2)⋯1 (with 0! = 1 by convention, so formulas don't break at the edges):
Order doesn't matter: combinations
Now: from 10 sprinters choose 3 for a relay squad — no medals, no order, just a group. Count the ordered selections (720) and notice each squad {A, B, C} was counted once per arrangement of itself: ABC, ACB, BAC, BCA, CAB, CBA — that's 3! = 6 repeats of the same group. So divide out the overcount:
Read (n over k) aloud as "n choose k" — the number of subsets of size k. Sanity-check the machinery on something listable: choose 2 from {A, B, C, D, E} → AB, AC, AD, AE, BC, BD, BE, CD, CE, DE — ten pairs, and indeed 5!/(2!3!) = 120/12 = 10. When a formula and a finger-count agree, you own the formula.
That one question — medals or squad? arrangement or subset? — is the entire decision. Devore's exercises (and exams) simply dress it in stories: PIN codes and seatings are permutations; committees, hands and samples are combinations.
Counting meets probability
Now combine with §2.2: draw 2 phones at random from a box of 8 where 3 are refurbished. What's the probability both are refurbished? All (8 choose 2) = 28 pairs are equally likely; the favourable pairs are those choosing 2 of the 3 refurbished, (3 choose 2) = 3. So P = 3/28 ≈ 0.107. One habit prevents the classic disaster: count numerator and denominator in the same mode — both unordered (as here) or both ordered, never mixed. The notes page's Worked Example 1 runs the full exam version of this pattern.
Check yourself 3a: how many 4-digit ATM PINs use four different digits? (Digits 0–9, order obviously matters.)
Ordered selection of 4 from 10: 10 · 9 · 8 · 7 = 5040. (All PINs with repeats allowed would be 10⁴ = 10,000 — the product rule with all stages at 10.)
Check yourself 3b: a badminton club has 8 players. How many doubles pairs can be formed? And if two of the 8 are siblings, what's the probability a randomly chosen pair is the siblings?
Pairs are unordered: (8 choose 2) = 8·7/2 = 28. Exactly one of those pairs is the siblings, so P = 1/28 ≈ 0.036. Numerator and denominator both counted as subsets — same mode.
Three patterns the tutorial sheet adds (same logic, new costumes)
Devore's §2.3 stops at permutations and combinations, but Tutorial 1 uses three more patterns. Good news: each one is the divide-out-the-overcount move you just learned for combinations, reworn.
Identical copies. How many arrangements of the letters of MISSISSIPPI? If all 11 letters were distinguishable: 11!. But the four S's are interchangeable — every arrangement got counted once per ordering of the S's among themselves (4! times), and likewise 4! for the I's and 2! for the P's. Divide the overcount out: 11!/(4! 4! 2!) = 34,650. Same move as dividing Pk,n by k!, just applied letter-group by letter-group.
Labelled groups. Split 10 officers into patrol (5), desk (2), reserve (3). Build it in stages with combinations: (10 choose 5)(5 choose 2)(3 choose 3) = 252 · 10 · 1 = 2520. Write those three fractions out and the denominators telescope into one clean formula: 10!/(5! 2! 3!) — the identical-copies formula again, because assigning people to labelled groups is the same as arranging the letters P,P,P,P,P,D,D,R,R,R.
The glue trick. When members of a group must sit together, glue each group into a block: arrange the blocks, then arrange within each block, product rule throughout. Four Americans, three French, three British in a row by nationality: 3! block orders × 4! · 3! · 3! internal orders = 6 · 864 = 5184.
Check yourself 3c: nine flags in a line — 4 identical white, 3 identical red, 2 identical blue. How many distinct signals? (This is Tutorial 1's first question.)
Identical copies: 9!/(4! 3! 2!) = 362880/288 = 1260. If you wrote 9!, you counted each signal 4! 3! 2! = 288 times over.
Check yourself 3d: three couples watch a film; each couple insists on sitting together in the row of six seats. How many seatings?
Glue each couple: 3! ways to order the three blocks, then 2! internal swaps per couple: 3! · 2! · 2! · 2! = 6 · 8 = 48.
4 · Conditional probability: updating on information §2.4
New information shrinks the world. A shipment of 200 phones, by brand and screen condition:
| Cracked | Fine | Total | |
|---|---|---|---|
| Brand A | 30 | 90 | 120 |
| Brand B | 10 | 70 | 80 |
| Total | 40 | 160 | 200 |
Pick a phone at random: P(cracked) = 40/200 = 0.20. Now someone tells you: it's a Brand A phone. The universe shrinks from 200 phones to the 120 in row A, and within that row, 30 are cracked: P(cracked | A) = 30/120 = 0.25. That vertical bar reads "given". Dividing top and bottom by 200 turns counts into probabilities and gives the general definition (for P(B) > 0):
Notice from the same table: P(A | cracked) = 30/40 = 0.75, while P(cracked | A) = 0.25. The two directions of conditioning are different questions with different answers. Keeping them straight is half of §2.4; the other half is two consequences of the definition:
The multiplication rule — read the definition backwards: P(A ∩ B) = P(A | B) · P(B). This is how you chain sequential events: P(first card ace and second card ace) = (4/52)(3/51) — the second factor already lives in the world where the first ace is gone. Trees make this mechanical: multiply along branches.
The law of total probability — when B can happen "via" several mutually exclusive routes A₁, …, Ak that cover everything (a partition), add the routes:
— a weighted average of the per-route rates, weighted by how likely each route is. It's the "overall defect rate of a factory with three production lines" formula, and the denominator of Bayes below.
Bayes' theorem: reversing the arrow
Often you know P(B | A) (easy direction, from data) but need P(A | B) (the question you actually care about). The honest way to feel it is with counts. A disease affects 1% of a population; the test catches 99% of the sick and false-alarms on 5% of the healthy. You test positive. How worried should you be? Take 10,000 people:
- Sick: 100. Of them, 99 test positive.
- Healthy: 9,900. Of them, 5% — that's 495 — also test positive.
- So the positive-test world contains 99 + 495 = 594 people, of whom only 99 are sick: P(sick | +) = 99/594 = 1/6 ≈ 0.17.
A 99%-accurate test, a positive result — and still only a one-in-six chance of disease, because the healthy crowd is so much bigger that even its small false-alarm rate out-produces the sick crowd's true alarms. That's the base-rate effect, and it is the single most examined (and most misunderstood-in-real-life) idea in the section. The formula version is just this computation in symbols — multiplication rule on top, total probability below:
On an exam, you may use the formula or literally draw the 10,000-people tree — both earn full marks if stated clearly. The notes page's Worked Example 2 is the standard three-production-lines version.
Three words that tutorial sheets and exams use for the pieces of this formula. The prior is P(sick) = 0.01 — what you believed before the test. The likelihood is P(+ | sick) = 0.99 (and P(+ | healthy) = 0.05) — how probable the data is under each possibility. The posterior is P(sick | +) = 1/6 — what you believe after. "The prior probabilities are 0.4 and 0.6; find the posterior" is asking for exactly this computation. One twist to expect: when the data is a compound event ("two of three tests detect"), the likelihoods aren't handed to you — section 5 shows how to build them, and Tutorial 2 Q5, from zero does the whole problem slowly.
Bayes theorem, the geometry of changing beliefs (3Blue1Brown, ~15 min) animates exactly the update-by-shrinking-worlds picture above. Watch it once, after this section — it's all on-syllabus for §2.4.
Check yourself: from the phone table, P(Brand B | cracked)? And why is it not equal to P(cracked | Brand B)?
Shrink to the cracked column: 40 phones, 10 of them Brand B → P(B | cracked) = 10/40 = 0.25. The other direction shrinks to row B: P(cracked | B) = 10/80 = 0.125. Different shrunken worlds, different denominators, different answers.
5 · Independence: when information is worthless §2.5
Sometimes the news that B happened tells you nothing about A: the shrunken world has the same proportion of A as the full one. That's independence:
The second form (multiply the probabilities) is the workhorse — and note it's a definition to check, not a vibe. Roll two dice: is A = "first die shows 6" independent of B = "sum is 7"? Feels like no — yet P(B | A) = P(second die is 1) = 1/6 = P(B). Independent! (Try the same with "sum is 12" and independence breaks: knowing the first die is 6 raises P(sum 12) from 1/36 to 1/6.) Intuition proposes; the multiplication check disposes.
Where independence genuinely comes from, in practice: physically unlinked mechanisms — separate coin tosses, separate components failing for separate reasons — or sampling with replacement. Sampling without replacement breaks independence (each draw changes the pool), though for tiny samples from huge populations it's approximately fine — a point Chapter 3 makes precise.
Mutually exclusive is not independence — it's the extreme opposite. If A and B are mutually exclusive (and both possible), then learning B happened tells you everything about A: it didn't. Check with the formula: P(A ∩ B) = 0 but P(A)P(B) > 0 — not equal, so dependent. Exams bait this every year.
The engineering payoff is system reliability. Components work or fail independently; systems combine them:
- Series (all must work): multiply the "works" probabilities. Two components at 0.9: 0.9 × 0.9 = 0.81 — series systems are weaker than their parts.
- Parallel (any one suffices): "at least one works" → complement trick: 1 − P(all fail) = 1 − 0.1 × 0.1 = 0.99 — redundancy buys reliability.
Mixed systems chain these two moves; the notes page's Worked Example 3 computes one in full.
Repeated trials: the sequence trick
Independence has one more everyday use: the same experiment repeated. A test detects an impurity with probability 0.8 each time it's run, runs being independent. Run it three times — what's the probability of detect, detect, no-detect, in that order? Multiply the three per-run probabilities: 0.8 × 0.8 × 0.2 = 0.128. (The third run happened and produced "no", an event of probability 0.2 — it earns its factor.) Now, the probability of exactly two detections in any order? List the sequences with two D's — DDN, DND, NDD — each has the same product 0.128 (order doesn't change a product), and no two of them can both happen, so add: 3 × 0.128 = 0.384. That's the whole method: multiply within a sequence, add across sequences. Module 2 will package it as the binomial formula C(n, k) pk(1 − p)n−k; for three or four trials, listing is faster and safer.
The same multiply-along-a-chain move prices "all different". Ten people, 365 equally likely birthdays: the second person differs from the first with probability 364/365; given that, the third differs from both with 363/365; … ; the tenth avoids the nine before with 356/365. Multiply the nine factors: 0.883. So "at least two share a birthday" is 1 − 0.883 = 0.117 — and a room needs only 23 people before that passes one half. Tutorial 2 Q6, from zero walks it through with a slider you can drag.
Check yourself: a component passes an independent inspection with probability 0.9. It is inspected four times. P(exactly three passes)?
Sequences with one fail among four: FPPP, PFPP, PPFP, PPPF — four of them (the fail can sit in any of 4 positions). Each: 0.9 × 0.9 × 0.9 × 0.1 = 0.0729. Add: 4 × 0.0729 = 0.2916. Audit: P(4 passes) = 0.6561, P(exactly 3) = 0.2916, P(exactly 2) = 6 × 0.0081 = 0.0486, P(exactly 1) = 4 × 0.0009 = 0.0036, P(0) = 0.0001 — total 1.0000 ✓.
Check yourself: toss a fair coin twice. A = "first toss H", B = "both tosses same". Independent?
P(A) = ½, P(B) = ½ ({HH, TT} out of four), P(A ∩ B) = P({HH}) = ¼ = ½ · ½. Yes — independent, even though B literally mentions the first toss. The formula outranks the feeling.
6 · You're ready — what to do next
That's §2.1–2.5 complete: the event dictionary, three axioms and their consequences, the order-matters/order-doesn't counting kit, conditioning as world-shrinking (with Bayes as the reversal), and independence as the multiply-check. Next:
- Work the examples: the notes page has the three exam-level worked examples (counting with defectives, three-lines Bayes, series-parallel reliability) and the practice table — counting gets deliberately more reps than anything else.
- Revise from the notes page, not this one.
Pin it to the syllabus and make it interactive:
"I'm studying Devore 9th ed. §2.3 and I never drilled permutations and combinations for JEE. Give me 6 short word problems that mix 'order matters' and 'order doesn't', one at a time — I answer which kind it is and compute; you correct me before moving on."
"Walk me through a Bayes' theorem problem (Devore §2.4 level) using the 10,000-people counting method AND the formula, then give me one to do myself with different numbers and check my answer."
One caution: AI answers can contain confident arithmetic errors — recompute any final fraction yourself.