MATH U113 · Probability & Statistics · Module 1
Probability
Devore (9th ed., Metric) §2.1–2.5 — sample spaces and events, axioms and properties, counting techniques, conditional probability, independence. The compressed revision map.
First time with this material? The Module 1 lesson teaches it slowly, including counting from absolute zero. Facts to hold onto, now confirmed by the handout: §2.1 and §2.3 are marked self-study — nobody in the hall gets a lecture on sample spaces or counting, so everyone teaches themselves (you have this page and the lesson; most classmates have only the book). The lectured part — §2.2, 2.4, 2.5 — lands in just two lectures (L1–2), so read ahead of them. The axioms-first formalism is new to everyone; the genuine JEE-drill advantage lives in exactly one section — counting, §2.3 — and the fix is reps, budgeted for in the practice table below.
Doubts filed on this module → the doubt clinic: king or ace and even first die or total 8 (§2.2 addition rule), 2 aces and 3 jacks (§2.3 combinations) — each worked from zero with the sample space drawn. Batch 2 (14 Sep): thirteen problem-set cards — Venn regions, hats, dinosaurs, phones, widgets, reliability, pairwise vs mutual, the circuit, vehicles, boxes, ball transfers, the director, and the two total-probability slides — filed by topic — see the clinic index (filter by 14 Sep) or go straight to sets, counting, independence, total probability. 15 Sep: six Devore worked examples she circled — 2.22, 2.26, 2.27–2.28, 2.29, 2.35, 2.36 — explained from zero on the conditional probability, counting and independence pages.
2.1 · Events are sets §2.1 self-study
Experiment → uncertain-outcome activity; sample space 𝒮 → the set of all possible outcomes; event → any subset of 𝒮, which occurs when the actual outcome lies in it. The dictionary: or = ∪, and = ∩, not = ′; A ∩ B = ∅ means mutually exclusive (disjoint) — they can't co-occur.
De Morgan's laws convert between "not-or" and "and-not" (draw a Venn to believe them): (A ∪ B)′ = A′ ∩ B′ — "neither" = "not this AND not that" — and (A ∩ B)′ = A′ ∪ B′. The first is used constantly: neither component failed, no defectives in the sample…
2.2 · Axioms and the rules they buy §2.2
Axioms: P(A) ≥ 0; P(𝒮) = 1; for mutually exclusive events, probabilities add. Interpretation: P(A) is the long-run relative frequency of A (the lesson's coin-toss widget shows the settling). Everything below follows:
- Complement rule: P(A′) = 1 − P(A) — the standard route to any "at least one" question: P(at least one) = 1 − P(none).
- P(∅) = 0; for any events, P(A) ≤ 1.
- Addition rule: P(A ∪ B) = P(A) + P(B) − P(A ∩ B) (the overlap was counted twice). Three events: add the singles, subtract the pairs, add back the triple.
- Equally likely outcomes (fair dice, shuffled decks, "at random" draws): P(A) = N(A)/N. School's "favourable over total" — valid only in the equally-likely case, and the reason counting (§2.3) matters at all.
2.3 · Counting techniques §2.3 self-study
One decision drives everything: is an outcome an arrangement (order matters) or a subset (order doesn't)?
| Task | Order? | Count | Signature stories |
|---|---|---|---|
| Multi-stage choice, stage i has ni options | — | n₁n₂⋯nk (product rule) | PINs with repeats allowed: 10⁴ |
| Arrange k of n distinct objects | Matters | Pk,n = n!(n−k)! = n(n−1)⋯(n−k+1) | medals, seatings, distinct-digit PINs |
| Choose a subset of k from n | Doesn't | (nk) = n!k!(n−k)! | committees, card hands, samples of items |
| Arrange n objects when some are identical copies (n₁ alike, …, nk alike) | Matters (copies interchangeable) | n!n₁! n₂! ⋯ nk! | flag signals, letters of MISSISSIPPI |
| Split n people into labelled groups of sizes n₁, …, nk | — | n!n₁! n₂! ⋯ nk! (same formula!) | 10 officers: 5 patrol, 2 desk, 3 reserve |
Tutorial 1 asks several counting questions that Devore §2.3 never names — the last two table rows, plus one trick. All are the same divide-out-the-overcount logic as the k! story above. Computed from the tutorial's own problems: 9 flags (4 white, 3 red, 2 blue alike) → 9!/(4! 3! 2!) = 1260 signals; 10 officers into patrol 5 / desk 2 / reserve 3 → 10!/(5! 2! 3!) = 2520 divisions. And the glue trick: when groups must sit together, glue each group into one block — arrange the blocks, then arrange inside each block. 4 Americans + 3 French + 3 British in national blocks → 3! · 4! · 3! · 3! = 5184. Fluency on all of these is what the counting drill builds. One more count Devore skips but the quiz used: derangements — orders in which nothing lands in its own place: D2 = 1, D3 = 2, D4 = 9 (of 4! = 24). Used in the matching problem, Module 2 worked example 2.
Unpack why the subset formula divides by k!
Count arrangements first (Pk,n), then notice each subset of size k was counted once per ordering of itself — k! times. Divide the overcount out: (n choose k) = Pk,n/k!.
Handy facts: 0! = 1; (n choose k) = (n choose n−k) (choosing who's in = choosing who's out); (n choose 0) = (n choose n) = 1. In probability use, the golden rule: numerator and denominator must count in the same mode — both ordered or both unordered.
Worked example 1Sampling defectives (the pattern behind half of Ch 2–3's exam questions)
A batch of 12 circuit boards contains 3 defective. An inspector draws 4 at random (without replacement). Find (a) P(exactly 1 defective), (b) P(at least 1 defective).
- Denominator — all equally likely samples, order irrelevant: (12 choose 4) = 495.
Unpack this step
(12 choose 4) = 12 · 11 · 10 · 94! = 1188024 = 495 — four falling factors over 4! = 24.
- (a) Favourable samples: build one by choosing 1 of the 3 defectives and 3 of the 9 good boards — product rule on two independent choices: (3 choose 1) · (9 choose 3) = 3 · 84 = 252.
Unpack this step
(9 choose 3) = 9 · 8 · 73! = 5046 = 84.
- (a) Probability: P = 252/495 = 28/55 ≈ 0.509. (Both counts are unordered — same mode ✓.)
- (b) Complement route — "at least 1" means avoid counting a pile of cases: P(none defective) = (9 choose 4)/(12 choose 4) = 126/495 = 14/55, so P(at least 1) = 1 − 14/55 = 41/55 ≈ 0.745.
- Sanity checks: both answers in [0, 1]; and (a) ≤ (b), as "exactly 1" is one way of having "at least 1". ✓
2.4 · Conditional probability, total probability, Bayes §2.4
Definition (for P(B) > 0): conditioning shrinks the sample space to B:
From a two-way table, P(A | B) is just "restrict to the B row/column and take the share" — the fastest exam method when counts are given. Chain sequential draws with the multiplication rule (each factor lives in the already-shrunk world): P(two aces) = (4/52)(3/51).
Law of total probability: if A₁, …, Ak partition 𝒮 (mutually exclusive, cover everything),
Bayes reverses the conditioning arrow: from the easy-to-measure P(effect | cause) to the wanted P(cause | effect). On exams, the tree/table layout below earns the same marks as the formula — use whichever is faster, but show the total-probability denominator explicitly.
Vocabulary the sheets use. Prior = P(Aj), your probability for a cause before the data; likelihood = P(B | Aj), how probable the data would be if that cause held; posterior = P(Aj | B), your probability after the data. Bayes turns priors and likelihoods into posteriors. When the data is a compound event — "exactly 2 of 3 tests detect", a run of results — the likelihoods are not given: build each P(B | Aj) first with the repeated-trials rule in §2.5, then run Bayes unchanged. Tutorial 2 Q5 is this pattern, worked from zero here.
Worked example 2Three production lines (total probability + Bayes)
A plant makes chips on three lines. Line A produces 50% of output with 4% defect rate, line B 30% at 2%, line C 20% at 10%. A chip is picked at random. (a) P(defective)? (b) Given it's defective, which line most likely made it?
| Route | P(line) | P(D | line) | Product |
|---|---|---|---|
| via A | 0.50 | 0.04 | 0.020 |
| via B | 0.30 | 0.02 | 0.006 |
| via C | 0.20 | 0.10 | 0.020 |
- (a) Total probability — the lines partition production, so add the routes: P(D) = 0.020 + 0.006 + 0.020 = 0.046 — a 4.6% overall defect rate.
- (b) Bayes — each route's share of the defective world: P(A | D) = 0.020/0.046 = 10/23 ≈ 0.435, P(B | D) = 0.006/0.046 = 3/23 ≈ 0.130, P(C | D) = 0.020/0.046 = 10/23 ≈ 0.435.
- Check: the three posteriors sum to 23/23 = 1 ✓ (they must — a defective chip came from some line).
- Interpret (exams ask): line C makes only 20% of the chips but owns ≈ 43.5% of the defects — tied with A, which makes half the output. High defect rate versus high volume: Bayes weighs both.
2.5 · Independence §2.5
- Definition / check: A, B independent ⟺ P(A ∩ B) = P(A)P(B) ⟺ P(A | B) = P(A). It's a computation to verify, not an intuition to assert.
- If A, B are independent, so are A, B′ — and A′, B′ (knowing nothing about B = knowing nothing about B′).
- Mutual independence of many events: multiplication holds for every sub-collection, not just pairs. In practice: physically separate mechanisms, or sampling with replacement.
- Pairwise is not mutual. Three events can pass all three pair checks and still fail the triple check. Standard example: two dice, A = red 3, B = green 4, C = total 7 — each pair meets in the single cell (3, 4), so P(A ∩ B) = 1/36 = (1/6)(1/6) ✓ for every pair, but P(A ∩ B ∩ C) = 1/36 ≠ 1/216: knowing A and B together makes C certain. Worked with the grid in the doubt clinic, Q10. Exam phrasing: "pairwise independent" asks for the three pair checks; "mutually independent" adds the triple.
- Reliability (independent components): series (all needed) → multiply "works"; parallel (one suffices) → 1 − Π(fails). Mixed systems: reduce subsystem by subsystem.
Worked example 3Series–parallel system reliability
Components 1 and 2 (each works with probability 0.9) are in parallel; that pair is in series with component 3 (works with probability 0.8). All fail independently. P(system works)?
- Reduce the parallel pair — it works unless both fail: P = 1 − (0.1)(0.1) = 0.99. (Complement trick: "at least one works" = 1 − "all fail" — legal because failures are independent.)
- Series with component 3 — both the (reduced) pair and component 3 must work; independence lets us multiply: P(system) = 0.99 × 0.8 = 0.792.
- Interpret: the redundant pair is near-perfect (0.99); the lone series component drags the system to 0.792 — a chain is as weak as its unduplicated link. This "spot the bottleneck" reading is a favourite short exam question.
Repeated independent trials — the sequence trick §2.5 → §3.4
- A specific sequence of results from independent trials has probability = the product of the per-trial probabilities. Three tests, detection probability 0.8 each, sequence DDN (detect, detect, miss): 0.8 × 0.8 × 0.2 = 0.128. Every trial that ran contributes a factor — including the ones that "failed".
- "Exactly k successes in n trials" is several sequences — the successes can occupy any C(n, k) of the positions — each with the same product, and mutually exclusive, so add: P = C(n, k) pk(1 − p)n−k. Exactly 2 detections in 3: 3 × 0.128 = 0.384. (This is the binomial pmf of §3.4 arriving early; for n ≤ 4, list the sequences rather than trust the formula.) Audit: the probabilities of all 2n sequences must add to 1.
- "All different" / "no match" chains. k people, N equally likely categories: P(all different) = (N/N) · ((N−1)/N) · ((N−2)/N) ⋯ ((N−k+1)/N) — each newcomer avoids the categories already taken (multiplication rule, one conditional factor per person); "at least two share" is the complement. Birthdays, N = 365: k = 10 gives 0.883 (share 0.117); the first k with share ≥ 0.5 is 23. Chain the fractions on a calculator; never compute 365k. Walk-through: Tutorial 2 Q6 from zero.
1) P(A ∪ B) = P(A) + P(B) without subtracting the overlap — only legal when disjoint. 2) "Favourable over total" on outcomes that aren't equally likely — check before you count. 3) Counting the numerator ordered and the denominator unordered (or vice versa) — pick one mode and hold it. 4) Swapping P(A | B) and P(B | A) — the disease-test example in the lesson shows they can differ by a factor of six. 5) "Mutually exclusive, so independent" — it's the opposite: disjoint events (both possible) are maximally dependent, since one occurring vetoes the other. 6) In Bayes, reporting the likelihood P(data | cause) as "the probability of the cause", or dropping the priors — the posterior needs both, and the arrow must be turned around. 7) "Two detections in three tests" priced as 0.8² — every trial that ran contributes a factor (×0.2 for the miss) and every ordering counts (×3).
Your minimal prerequisite kit for this module
- Fraction arithmetic and simplification — most answers land as fractions like 41/55; simplify by common factors.
- Percent ↔ decimal ↔ fraction, instantly (0.046 = 4.6%).
- Set/Venn vocabulary: subset, union, intersection, complement — 9th-class sets chapter is plenty.
- Factorials: n! grows fast; cancel before multiplying (12!/8! = 12·11·10·9, never compute 12! outright).
- Systematic listing: for small cases, write outcomes in a fixed order (tree discipline) — the checking tool for every counting formula.
- Running products on a calculator: chain fraction × fraction × … and read the running value after each step (it must stay in [0, 1] and only fall); never compute a huge power like 36510 and divide.
What to practise in Devore
| Skill | Where | How many |
|---|---|---|
| Event algebra, Venn diagrams, mutually exclusive or not | §2.1 exercises | 3–4 |
| Addition & complement rules; equally-likely computations | §2.2 exercises | 4–5 |
| Counting — the catch-up section: extra reps here, mixing order/no-order until the decision is automatic | §2.3 exercises + the counting drill (covers the tutorial's beyond-Devore patterns) | 8–10 + one drill sheet/day for a week |
| Conditional probability from tables & trees; multiplication rule | §2.4 exercises | 4–5 |
| Total probability & Bayes (state the denominator!) | §2.4 exercises | 3–4 |
| Independence checks; series/parallel reliability | §2.5 exercises | 4–5 |
Prefer odd-numbered exercises (answers to selected odd ones are in the back). Metric Version numbering may differ from the US edition — choose by section and skill, not by numbers copied from elsewhere.