MATH U113 · Probability & Statistics · R lab support
R lab 2 · Simulating distributions
The lab sheet "Simulation of Probability Distributions in R": binomial §3.4, Poisson §3.6, normal §4.3, exponential §4.4. It covers the four R function families, the three traps that cost marks, past mid-sem answers checked in one line of R, and the sheet's mistakes corrected.
What this is for: the 10% Lab Exam in R (open book, date not yet announced). The mid-sem on 7 Oct is closed book with no R, so don't spend mid-sem hours here beyond one side benefit (§3): R can check your hand answers to past papers in one line.
Where you actually stand: every distribution on the sheet (binomial, Poisson, exponential, normal) is one you've already met in Modules 2–3, with the same means and variances. The genuinely new things are R's naming scheme (§1) and simulation itself (§4). Those are new to everyone in the hall; no school syllabus, JEE or otherwise, teaches rbinom. The sheet also previews two post-midsem ideas, the Law of Large Numbers and the normal approximation, which get lectures in Module 5 §5.3–5.4.
1 · One naming scheme, four jobs
R names every distribution function with a one-letter prefix (what you want) plus the distribution's short name (which one). Learn the four prefixes once and you know forty functions.
| Prefix | Answers the question | Devore's name for it | Example |
|---|---|---|---|
| d | "How high is the pmf/pdf at x?" | Discrete: p(x) = P(X = x). Continuous: f(x), a height, not a probability | dbinom(3, 10, 0.4) |
| p | "What's the probability of X ≤ x?" | The cdf F(x); for the normal, Φ((x − μ)/σ); for Devore's binomial table, B(x; n, p) | pnorm(55, 50, 5) |
| q | "Which x has this much probability below it?" | The percentile, the inverse of the cdf (Devore's η(p), zα) | qnorm(0.95) |
| r | "Give me N random values from it." | (No hand equivalent. This is simulation.) | rexp(10000, rate = 2) |
The four distributions on the sheet, with the argument names R expects. The last column is where the marks go:
| Distribution | R short name + arguments | Mean · Variance | R's second parameter is… |
|---|---|---|---|
| Binomial(n, p) §3.4 | binom(x, size = n, prob = p) | np · np(1 − p) | — |
| Poisson(μ) §3.6 | pois(x, lambda = μ) | μ · μ | — (the sheet writes λ; Devore writes μ) |
| Exponential(λ) §4.4 | exp(x, rate = λ) | 1/λ · 1/λ2 | the rate, not the mean |
| Normal(μ, σ2) §4.3 | norm(x, mean = μ, sd = σ) | μ · σ2 | the standard deviation, not the variance |
Read the picture as one sentence: pnorm(55, 50, 5) = 0.8413 is the shaded area (the probability of landing below 55), and qnorm(0.8413, 50, 5) = 55 undoes it. dnorm(55, 50, 5) = 0.0484 is only the curve's height at 55. For a continuous variable P(X = 55) = 0, so a d value is never an answer to "what's the probability" there (Module 3's "a single point has zero area").
Check yourself: which one call gives the 95th percentile of the sheet's normal, N(50, 5²)? Roughly what number should it print?
qnorm(0.95, mean = 50, sd = 5). It is a percentile, so it's a q. By hand, z0.05 = 1.645, so the answer is 50 + 1.645 × 5 = 58.22.
2 · The three traps
All three produce a plausible-looking number with no error message, and that's why they cost marks. Each is a mismatch between how Devore writes a distribution and how R wants it.
Trap 1 · rnorm/pnorm take the standard deviation
Devore writes X ~ N(80, 100) with the variance second. R's third argument is sd, so pass √100 = 10. The lab sheet itself switches between "Normal(μ, σ2)" and "Normal(μ, σ)" in different places. Always name the argument (sd = 10) and the slip becomes visible.
Trap 2 · rexp/pexp take the rate
Past papers describe lifetimes by their mean ("mean lifetime 12 000 hours"). The rate is the reciprocal: λ = 1/12000. Writing pexp(1500, 12000) silently uses a mean of 1/12000 hours, and the answer comes out as 1.
Trap 3 · p means "at most", including the endpoint
For a discrete variable, pbinom(k, n, p) is P(X ≤ k), and it includes k. So "at least k" is everything except 0, …, k − 1:
pbinom(3, 10, 0.4) adds the shaded bars 0–3. "At least 4" is the unshaded rest.This is the same boundary habit as continuity correction. For continuous distributions it disappears: P(X ≥ a) = 1 − F(a) exactly, because the single point a carries no probability.
Check yourself: X ~ Poisson(4). Write the R call for P(X > 7), then say what the tempting wrong call is.
"More than 7" means 8 or more, so the call is 1 - ppois(7, lambda = 4) = 1 − 0.9489 = 0.0511. The tempting wrong call is 1 - ppois(8, 4), which leaves out X = 8. The trap is deciding whether the boundary value belongs to the event. Translate the words into a set of integers first ({8, 9, …}), then write the complement.
3 · R as a checker: past mid-sem answers in one line
Each card solves a real past-paper question by hand first, as the exam needs, then checks it in R. That's the only way R helps before 7 Oct. It's also exactly the skill the open-book lab exam tests.
Worked example 1Binomial tail: exact vs. normal approximation 2026 mid-sem Q3
X ~ Bin(1000, ½). Find P(X ≥ 525) (full question).
- By hand (the exam route): μ = np = 500 and σ = √(np(1 − p)) = √250 = 15.81. With the continuity correction, "≥ 525" becomes "> 524.5", so z = (524.5 − 500)/15.81 = 1.55 and 1 − Φ(1.55) = 1 − 0.9394 = 0.0606.
- The normal approximation in R, a direct copy of step 1:
1 - pnorm(524.5, mean = 500, sd = sqrt(250))→ 0.06063. - The exact answer, which you can't do by hand:
1 - pbinom(524, 1000, 0.5)→ 0.06061. The approximation was accurate to 4 decimal places, and the continuity correction is why. - The two traps, measured:
1 - pbinom(525, 1000, 0.5)(trap 3) gives 0.05337. Dropping the continuity correction,1 - pnorm(525, 500, sqrt(250)), gives 0.05692. Both are wrong in the second decimal place.
Worked example 2Normal: the sd-vs-variance trap in action 2025 mid-sem Q1
X ~ N(80, σ2) with σ = 10 (found in the question's first part: full question). Find P(X ≤ 60) and P(65 < X < 75).
- By hand: P(X ≤ 60) = Φ((60 − 80)/10) = Φ(−2) = 0.0228. For the interval, Φ(−0.5) − Φ(−1.5) = 0.3085 − 0.0668 = 0.2417.
- In R:
pnorm(60, mean = 80, sd = 10)→ 0.02275. For the interval, subtract twopvalues:pnorm(75, 80, 10) - pnorm(65, 80, 10)→ 0.24173. - Trap 1, measured:
pnorm(60, 80, 100), with the variance passed as sd, gives 0.4207, eighteen times too big. No normal distribution puts 42% of its mass two or more standard deviations below its mean (the true figure is always 2.28%), so that number should set off an alarm.
Worked example 3Poisson inside a binomial 2026 mid-sem Q5
Defects occur at Poisson(2) per m². (a) P(at most 1 defect in a region). (b) Of 5 independent regions, P(exactly 3 have no defect) (full question).
- (a) by hand: p(0) + p(1) = e−2 + 2e−2 = 3e−2 = 0.4060. In R:
ppois(1, lambda = 2). "At most 1" is exactly whatpmeans, so no adjustment is needed. - (b) by hand: a region is empty with probability q = e−2 = 0.1353. The number of empty regions out of 5 is Bin(5, q), so P = C(5,3) q3(1 − q)2 = 10 × 0.002479 × 0.7477 = 0.0185.
- In R, nesting one
dinside another:dbinom(3, size = 5, prob = dpois(0, 2))→ 0.01853. The R line has the same two-stage structure as the hand solution. Writing it is a good test of whether you understood the structure.
Worked example 4Exponential from a mean 2026 mid-sem Q4, first step
A component's lifetime is exponential with mean 12 000 hours. Find P(it fails within 1 500 hours) (full question).
- Convert the mean to a rate: λ = 1/12000 per hour.
- By hand: F(1500) = 1 − e−1500/12000 = 1 − e−0.125 = 0.1175.
- In R:
pexp(1500, rate = 1/12000)→ 0.11750. The trap-2 version,pexp(1500, 12000), prints 1, which should be an instant red flag.
4 · What simulation shows you
The sheet's recipe is the same for all four distributions:
set.seed(123) # 1. same "random" numbers every run
x <- rbinom(10000, size = 10, prob = 0.4) # 2. draw N values
mean(x); var(x) # 3. compare with np = 4, np(1-p) = 2.4
hist(x, breaks = seq(-0.5, 10.5, 1), probability = TRUE)
points(0:10, dbinom(0:10, 10, 0.4), col = "red", pch = 19) # 4. overlay theory
Two details carry the whole idea. First, probability = TRUE puts the histogram on the density scale (bar area = proportion). That is the unequal-width rule from Lab 1, and it's what makes the bars comparable to a pmf or pdf at all. Second, breaks = seq(-0.5, 10.5, 1) centres one bar of width 1 on each integer. Its height then equals the fraction of draws at that value, which should match dbinom there.
Now try it. The widget below simulates the same four distributions in your browser. It won't reproduce R's exact numbers, because it uses a different random generator, but it follows the same recipe. Slide N up and watch the bars settle onto the red theory while the errors in the readout shrink.
What you should see (this is the sheet's Exercise (a)). At N = 100 the mean is often off by a few percent and the histogram is ragged. At N = 10 000 it's typically within about 1%, and for the normal within a tenth of that. The improvement is steady but slow: 100 times more draws buys only 10 times more accuracy. The typical error of a sample mean is σ/√N, and √100 = 10.
Unpack: why σ/√N (a Module 5 preview)
The sample mean X̄ of N independent draws has variance σ2/N, so its standard deviation is σ/√N. That's proved in §5.4 after the mid-sem. For now, just recognise the pattern: quadrupling N halves the error.
Check yourself: for the sheet's Bin(10, 0.4) at N = 100, how big is the typical error of the sample mean, as a percentage of the true mean 4?
σ = √2.4 = 1.549, so the typical error is 1.549/√100 = 0.155, which is 0.155/4 ≈ 3.9%. That's why N = 100 looks ragged in the widget. At N = 10 000 it drops to 0.39%, which matches the sheet's own 3.9849 against 4.
The sheet's other exercises, briefly. (b) Bin(10, 0.5) is exactly symmetric about 5, because p(x) = p(10 − x) when p = ½. (c) The fraction of normal draws within μ ± σ should come out near Φ(1) − Φ(−1) = 0.6827, and within ±2σ near 0.9545. mean(x > a & x < b) works because R counts TRUE as 1, so the mean of a TRUE/FALSE vector is a proportion. (d) Poisson(0.5) is heavily L-shaped, with P(0) = e−0.5 = 0.607. Poisson(20) looks close to a bell. That's the normal approximation to the Poisson, which you'll meet formally in Module 5.
5 · Corrections to the lab sheet
The formulas, the code and the theoretical values on the sheet are all correct. Four of its "What to Observe" remarks are not, and the notation slips once. If your own run disagrees with these particular sentences, you haven't done anything wrong.
| Sheet says | Actually | Why |
|---|---|---|
| Bin(10, 0.4): "roughly 96% of the simulated mass falls between x = 1 and x = 7" | 98.2% | The mass outside is p(0) + p(8) + p(9) + p(10) = 0.0060 + 0.0106 + 0.0016 + 0.0001 = 0.0183 |
| "Across all four distributions … agree … to within roughly 0.5% at N = 10 000" | Not for the Poisson variance: 3.9061 vs 4 is 2.3% off | Sample variances wobble more than sample means |
| Exercise (a): at N = 100 000 "the discrepancy typically shrinks to well under 0.1%" | Typical error of the mean: 0.12% (binomial), 0.16% (Poisson), 0.32% (exponential), 0.03% (normal) | Relative error = (σ/√N)/μ. For the exponential, σ = μ, so it's 1/√100000 = 0.32% |
| Exponential histogram "shows the characteristic 'memoryless' decaying shape" | The shape is a decaying curve. Memoryless is a different fact | Memorylessness, P(X > s + t | X > s) = P(X > t), is a property of probabilities. You can't see it by looking at a histogram |
| Summary table: "Normal(μ, σ)" | Devore and this course write N(μ, σ2) | R's argument is the sd. The textbook's parameter is the variance. See trap 1 |
One more caveat: the sheet's "Expected Output" numbers come from set.seed(123). They should match on current R versions if you run each block exactly as printed, starting from its own set.seed. If you skip a set.seed line or run blocks out of order, the numbers will differ. That's harmless.
1) Variance where R wants the sd in rnorm/pnorm/qnorm. 2) Mean where R wants the rate in rexp/pexp. 3) "At least k" written as 1 - pbinom(k, …). It must be k - 1 for discrete distributions. 4) Reporting a d value as a probability for a continuous distribution. dnorm is a height, and it can even exceed 1 (look at the exponential with rate 2 near 0). 5) A histogram without probability = TRUE under a pmf/pdf overlay. The bars are then counts in the thousands and the theory curve sits flat along the bottom.
Your minimal prerequisite kit for this lab
- The four pmfs/pdfs and their mean/variance: binomial and Poisson (Module 2), exponential and normal (Module 3). The second table in §1 above is the whole list.
- Complement rule P(A′) = 1 − P(A), plus the habit of writing "at least k" as an explicit set of integers.
- Standardising z = (x − μ)/σ and reading Φ, to check every
pnormby hand. - Lab 1's R basics: vectors,
mean/var,hist(probability = TRUE)(Lab 1 page, §1.5).
What to practise
| Skill | Where | How many |
|---|---|---|
| Run the sheet's four blocks, then change one parameter each (e.g. p = 0.5, λ = 0.5 and 20). These are Exercises (b) and (d) | Lab 2 sheet §1–4, §6 | 4 runs |
Check hand answers with d/p calls: binomial and Poisson | Devore §3.4, §3.6 odd exercises (answers in the back) | 4–5 |
Check hand answers with pnorm/qnorm/pexp, naming sd = and rate = every time | Devore §4.3, §4.4 odd exercises | 4–5 |
| Re-derive this page's four worked examples in R without looking | §3 above | 4 |
Budget: about 90 minutes total, after the 7 Oct mid-sem unless the lab exam gets announced first.
Pin it to the course and to base R so it doesn't reach for packages or theory you don't have:
"I'm learning base R for a probability course (Devore 9th ed., §3.4, §3.6, §4.3, §4.4). Give me 6 short word problems on the binomial, Poisson, exponential and normal. Some should say 'at least', some give a mean lifetime, and some give a variance. For each, I'll write the one-line R call. Mark my call and tell me if I fell into the sd-vs-variance, rate-vs-mean or pbinom(k − 1) trap."
"Explain, without calculus beyond Devore ch. 4, why hist(x, probability = TRUE) and dnorm() can be drawn on the same axes but hist(x) and dnorm() can't."
Cross-check any number it gives you by running the R line yourself. That's the point of the lab.