MATH U113 · Previous-year Quiz 1 · rebuilt from the ground up
The two quiz questions, from zero
Last year's Quiz 1 had exactly two questions, and this year's is drawn from the same two ideas: a binomial count inside a Bayes problem, and expectation and variance of a function of a random variable. This page rebuilds both from first principles — every symbol named before it is used, every arithmetic step shown. The compressed version, if you only want the answer, is in the tutorial companion.
These are not hard questions hiding a trick. They are two standard patterns with fiddly arithmetic, under a clock, with no partial credit. That combination — not the difficulty — is what costs marks. So this page spends its time on the two places people actually lose them: reading which probability is being asked for, and getting the decimals right.
Everything here is §2.4–2.5 (conditional probability, Bayes) and §3.1–3.4 (pmf, expectation, variance, binomial) — all of it taught, all of it in Modules 1 and 2. Nothing on this page is new syllabus.
The format, and what it rewards
From last year's paper, which is in the repo: 20 minutes, 20 marks, closed book, two questions of 5 + 5. You write only a final answer in a blank, rounded to five decimal places. Rough work is not evaluated, and overwriting voids the answer. About four numeric variants circulate in the room — same structure, shuffled numbers — so a memorised answer is worth nothing and a memorised method is worth everything.
1. Sanity-check every probability before you write it. Anything outside [0, 1] is wrong, and a posterior that moved the wrong way is wrong. Two seconds each.
2. Know the rounding cold. Five decimal places means five digits after the point, including trailing zeros: 0.471803… is written 0.47180, not 0.4718. Writing four digits where five are asked is a real way to lose a whole answer.
3. Do the arithmetic in rough space, check it once, then write once. Overwriting voids it.
Q1 · Two coins, five tosses §2.5 + §3.4
A box contains two coins C₁ and C₂. The probability of choosing coin C₁ is 0.4. The probability of getting a head when C₁ is tossed is 0.4; for C₂ it is 0.7. One coin is chosen at random and tossed 5 times.
(i) Find the probability of observing at least 3 heads given that coin C₁ was selected. [5]
(ii) Find the probability that the selected coin was C₂ given exactly 3 heads were observed. [5]
Step 0Name everything before computing anything
Three sentences, three symbols. Write them down — on a 20-minute paper this takes fifteen seconds and prevents the one mistake that costs five marks:
And let X = the number of heads in the 5 tosses. Notice that P(C₂) is not given directly — you get it from "the other coin", because exactly one of the two is chosen. That is the first small step people skip.
Part (i) — the word "given" does all the work
Step 1Read what is being conditioned on, and delete the rest
The question says "given that coin C₁ was selected". That means: stop imagining the box. The coin is settled. You are in a world where the only coin is C₁, and it lands heads with probability 0.4 every toss.
P(C₁) = 0.4 plays no part in part (i). It is given information about a step that has already happened. Multiplying your answer by 0.4 at the end is the most common way this question is failed, and it is tempting precisely because the number 0.4 appears twice in the problem for two completely different reasons — once as the chance of picking C₁, once as the chance of a head with C₁. They are different quantities that happen to share a value in this variant. In variant C they don't (0.4 and 0.6), which is exactly why the paper-setter shuffles them.
Step 2Why the head-count is binomial — check the four conditions
"Binomial" is not a word to pattern-match; it is four conditions, and you should be able to tick them off:
- Fixed number of trials. Exactly 5 tosses — decided in advance. ✓
- Two outcomes per trial. Head or tail. ✓
- Same probability every trial. The coin doesn't change between tosses: 0.4 each time. ✓
- Independent trials. One toss tells you nothing about the next. ✓
All four hold, so X ~ Binomial(n = 5, p = 0.4), and
Unpack this step — where the C(5, k) comes from
The probability of one particular sequence with 3 heads — say HHHTT — is (0.4)(0.4)(0.4)(0.6)(0.6) = (0.4)³(0.6)², because the tosses are independent so probabilities multiply. But HHHTT is not the only way to get 3 heads: HHTHT, HTHHT, THHHT and so on all work, and each has the same probability. So you count the arrangements and multiply. The number of ways to choose which 3 of the 5 positions are heads is C(5,3) = 10. That is the whole content of the formula: one sequence's probability, times how many such sequences there are.
The three values you need here, worth memorising as a row of Pascal's triangle: C(5,3) = 10, C(5,4) = 5, C(5,5) = 1.
Step 3"At least 3" means add the three top terms
There is no single formula for "at least 3". X can only be 0, 1, 2, 3, 4 or 5, so "at least 3" is the three cases 3, 4 and 5, and since they can't happen together you add:
Now grind it out, one term at a time, writing each down:
| k | the term | value |
|---|---|---|
| 3 | 10 × (0.4)³ × (0.6)² = 10 × 0.064 × 0.36 | 0.23040 |
| 4 | 5 × (0.4)⁴ × (0.6)¹ = 5 × 0.0256 × 0.6 | 0.07680 |
| 5 | 1 × (0.4)⁵ = 0.01024 | 0.01024 |
| — | total | 0.31744 |
Answer (i): 0.31744 — already exactly five decimal places.
Step 4Check it before you write it
- In range? 0.31744 is between 0 and 1. ✓
- Does the size make sense? A coin that comes up heads only 40% of the time should give you 3+ heads out of 5 less than half the time. 0.317 is comfortably below 0.5. ✓ If you had got 0.68, you would have used p = 0.6 by mistake — which is exactly what variant C asks for, so the check matters.
- Faster alternative? Only if you prefer: P(X ≥ 3) = 1 − P(0) − P(1) − P(2). Same work, three terms either way. Don't bother switching.
Part (ii) — the direction flip, which is what Bayes is
Step 5Notice that the question has been turned around
Everything you were given runs one way: coin → heads. You know P(3 heads | C₂). What you are asked runs the other way: heads → coin. You want P(C₂ | 3 heads).
P(A | B) and P(B | A) are different numbers, and Bayes' theorem is the machine that converts one into the other. Every Bayes problem in this course — the impurity test in Tutorial 2, the disease-test examples in Module 1, this coin — is that same flip. If you can spot "I was told it one way and asked it the other way", you have identified the method, and the rest is arithmetic.
Three words the textbook uses, because the question is easier once they are named:
- Prior — what you believed before seeing data: P(C₂) = 0.6.
- Likelihood — how well each hypothesis explains the data you saw: P(3 heads | C₂) = 0.30870.
- Posterior — the updated belief, which is what's asked: P(C₂ | 3 heads).
Step 6Weigh each explanation: prior × likelihood
The evidence — exactly 3 heads — could have come about in two ways: through C₁, or through C₂. Give each route a weight equal to how likely you were to take that route times how likely that route was to produce this evidence:
| Route | prior | likelihood: P(exactly 3 heads) | weight |
|---|---|---|---|
| through C₁ | 0.4 | 10 (0.4)³(0.6)² = 0.23040 | 0.09216 |
| through C₂ | 0.6 | 10 (0.7)³(0.3)² = 0.30870 | 0.18522 |
Note that the likelihoods use exactly 3 heads — a single binomial term each, not the "at least 3" sum from part (i). Reusing part (i)'s 0.31744 here is a real and common slip; the two parts ask different things.
Unpack this step — computing 10 (0.7)³(0.3)²
(0.7)³ = 0.343. (0.3)² = 0.09. Then 0.343 × 0.09 = 0.03087, and ×10 = 0.30870. Keep all the digits as you go; rounding at an intermediate step is how a correct method produces a wrong fifth decimal.
Step 7The posterior is C₂'s share of the total weight
The two weights add up to the total probability of seeing exactly 3 heads at all (that is the law of total probability). The answer is simply how much of that total belongs to C₂:
Answer (ii): 0.66775 (the unrounded value is 0.667748…, so the fifth decimal rounds 4→ stays, giving 0.66775).
C₂ started as the more likely coin (prior 0.6) and it explains 3 heads better than C₁ does (0.30870 vs 0.23040). Both pull the same way, so the posterior must be above 0.6. It is: 0.66775. If you had written 0.33225 — the complement — this check catches it instantly. Always ask: did the evidence push my belief in the direction it should have?
Q2 · A function of a two-toss count §3.2–3.3
Consider tossing a fair coin twice. Let X denote the total number of heads after the two tosses. Then
(i) E[(X + 1)²] is ______ [5]
(ii) If V[aX² + 1] = 2.25, find the positive value of a. [5]
Step 1Build the pmf yourself — this is the step everything else stands on
Nothing can be computed until you know what values X takes and with what probabilities. Don't reach for a formula; list the outcomes. Two tosses of a fair coin give four equally likely results, each with probability ¼:
Check they sum to 1: ¼ + ½ + ¼ = 1 ✓. This is the same "construct it from the sample space" routine as the dice-grid walkthrough in the Module 2 lesson — and it is worth being fast at, because every part of this question is a sum over these three rows.
Unpack this step — you could also say X ~ Binomial(2, ½)
You can: P(X=1) = C(2,1)(½)¹(½)¹ = 2 × ¼ = ½, and so on — same three numbers. Listing the four outcomes is quicker here and far less error-prone under time pressure. Use the formula when n is too big to list, as in Q1.
Part (i) — expectation of a function, term by term
Step 2You do not need the distribution of (X + 1)²
This is the idea the question is really testing. To average a function of X, you do not have to work out what values (X+1)² takes and how likely each is. You take each value X can be, apply the function to it, and weight by X's own probability:
Three rows, three multiplications. Write the table out — it is faster than being clever and it is checkable:
| x | p(x) | (x + 1)² | product |
|---|---|---|---|
| 0 | ¼ | 1² = 1 | 0.25 |
| 1 | ½ | 2² = 4 | 2.00 |
| 2 | ¼ | 3² = 9 | 2.25 |
| — | — | total | 4.50 |
Answer (i): 4.5 — write it as 4.50000 if the paper asks for five decimals.
Step 3Cross-check by expanding — thirty seconds, and it catches slips
Expectation is linear, which means it passes through sums and constants: E[X² + 2X + 1] = E[X²] + 2E[X] + 1. So compute the two moments once and reuse them:
| working | value | |
|---|---|---|
| E[X] | 0(¼) + 1(½) + 2(¼) | 1 |
| E[X²] | 0(¼) + 1(½) + 4(¼) | 1.5 |
Both routes give 4.5. Keep E[X] = 1 and E[X²] = 1.5 written down — part (ii) needs them, and recomputing wastes clock.
Here E[X²] = 1.5 but (E[X])² = 1. They are never equal unless the variable is a constant, and the gap between them is the variance: V[X] = E[X²] − (E[X])² = 1.5 − 1 = 0.5. Squaring the average is not the average of the squares.
Part (ii) — the variance of a squared variable
Step 4Give X² a name and it stops being frightening
The expression V[aX² + 1] looks like it needs new machinery. It doesn't. Let Y = X². Then Y is just another random variable, with its own three values — and its probabilities are inherited unchanged from X, because squaring doesn't move any probability around, it only relabels the values:
The question now reads V[aY + 1] = 2.25, which is a shape you already know.
Step 5Two rules for what scaling and shifting do to variance
Variance measures spread. So:
- Adding a constant does nothing. V[Y + 1] = V[Y] — sliding every value up by 1 moves the whole distribution but doesn't stretch it. The +1 in the question is there purely to see whether you know this.
- Multiplying by a constant scales the spread — and variance is in squared units, so the constant comes out squared. V[aY] = a²V[Y].
Unpack this step — why a squares but the shift vanishes
Variance is the average of (value − mean)². Add 1 to every value and the mean also rises by 1, so every gap (value − mean) is unchanged — variance unchanged. Multiply every value by a and the mean also multiplies by a, so every gap multiplies by a; but the gaps are squared before averaging, so the variance multiplies by a². The sign of a disappears in the squaring, which is why the question has to ask for the positive value.
Step 6Find V[Y], then solve for a
Use the same variance formula on Y, reading its values 0, 1, 4 straight off the figure:
| working | value | |
|---|---|---|
| E[Y] | 0(¼) + 1(½) + 4(¼) | 1.5 |
| E[Y²] | 0²(¼) + 1²(½) + 4²(¼) = ½ + 4 | 4.5 |
| V[Y] | E[Y²] − (E[Y])² = 4.5 − 1.5²= 4.5 − 2.25 | 2.25 |
(E[Y] = E[X²] = 1.5 is the number you already had from part (i) — that is why it was worth keeping.)
Answer (ii): a = 1 (positive root, as asked).
V[Y] = 2.25 is fixed by the coin — it does not depend on the variant. So in every version of this paper, a² = (the given variance) ÷ 2.25, and the numbers are chosen so that a comes out a whole number: the four variants give a = 1, 2, 3, 4. If your a is not a tidy integer, you have made an arithmetic slip — go back and check E[Y²] first, since 16 × ¼ is where the slips happen.
All four variants, worked
Roughly four papers circulate. Same two questions, shuffled numbers. Use this table to practise the method four times, not to memorise answers — and note that Q1(i) depends only on C₁'s head-probability, so it takes just two distinct values across the whole room.
| Variant | P(C₁) | P(H|C₁) | Q1 (i) P(X ≥ 3 | C₁) | Q1 (ii) P(C₂ | 3H) | Q2 (i) | Q2 (i) value | Q2 (ii) | Q2 (ii) a |
|---|---|---|---|---|---|---|---|---|
| A | 0.4 | 0.4 | 0.31744 | 0.66775 | E[(X+1)²] | 4.5 | V[aX²+1] = 2.25 | 1 |
| B | 0.6 | 0.4 | 0.31744 | 0.47180 | E[(X−1)²] | 0.5 | V[aX²+2] = 9 | 2 |
| C | 0.4 | 0.6 | 0.68256 | 0.57262 | E[(X+2)²] | 9.5 | V[aX²+3] = 20.25 | 3 |
| D | 0.6 | 0.6 | 0.68256 | 0.37323 | E[(X−2)²] | 1.5 | V[aX²+4] = 36 | 4 |
Two patterns worth seeing in that table. In Q1(ii), the posterior falls as P(C₁) rises — more prior weight on C₁ means the same evidence leaves less belief on C₂; that is the sanity check from Step 7, visible as a column. And in Q2(i), E[(X+c)²] = E[X²] + 2cE[X] + c² = 1.5 + 2c + c² — one formula, four values of c.
Check yourself — do variant C from scratch, then open this
(i) p = 0.6: 10(0.6)³(0.4)² + 5(0.6)⁴(0.4) + (0.6)⁵ = 0.34560 + 0.25920 + 0.07776 = 0.68256.
(ii) branches: C₁ → 0.4 × 10(0.6)³(0.4)² = 0.4 × 0.34560 = 0.13824; C₂ → 0.6 × 0.30870 = 0.18522. Posterior = 0.18522 / 0.32346 = 0.57262. Sanity: still above the prior 0.6? No — 0.573 < 0.6, and that is correct here, because with p = 0.6 coin C₁ now explains 3 heads almost as well as C₂ does, so the evidence pulls belief slightly away from C₂. Notice the check is not "always goes up" — it is "moves the way the likelihoods say".
Q2: E[(X+2)²] = 1.5 + 4 + 4 = 9.5; a² = 20.25/2.25 = 9, so a = 3.
The 20-minute plan
- Minute 0–1: write the givens as symbols, including the one you have to infer (P(C₂) = 1 − P(C₁)). Read whether part (i) says "at least" or "exactly".
- Minutes 1–7: Q1. Part (i) is three binomial terms added. Part (ii) is two branch weights and a ratio. Write every intermediate number; don't round until the end.
- Minutes 7–13: Q2. Build the three-row pmf first, always. Get E[X] and E[X²] once and reuse them in both parts.
- Minutes 13–17: check. Every probability in [0, 1]; the posterior moved the way the likelihoods say; a came out a whole number; part (i) of Q1 didn't get multiplied by a prior.
- Minutes 17–20: transcribe once, to five decimal places, trailing zeros included. Write each answer exactly once — overwriting voids it.
- "Give me a two-coin Bayes problem with a binomial likelihood, in the style of Devore §2.5 with n = 5 tosses. Don't solve it. After I answer, tell me only whether my prior, my likelihood and my final ratio are each right."
- "I get P(C₂ | 3 heads) = [my number]. Without giving me the answer, tell me whether it should be larger or smaller than the prior P(C₂), and why — in terms of which coin explains 3 heads better."
- "Set me four quick drills on V[aX + b] and E[g(X)] for a random variable taking three values, Devore §3.3. Ask one at a time and wait for my answer."
Don't ask it to work the quiz question for you — the arithmetic is the skill being tested, and the paper gives no credit for a method you can't execute under a clock.