Distribution of the Sample Mean

A2 · S2 · 13 min

Take a random sample and work out its mean. Take another sample of the same size and you get a different mean. Before the sample is drawn, the sample mean is uncertain, so it is a random variable with its own distribution. This note finds that distribution: its mean, its variance, and its exact shape when the population is normal. These facts drive every confidence interval and every test for a mean on Paper 6, and questions that ask directly for P(Xˉ>k)P(\bar{X} > k) appear on most papers.

The sample mean is a random variable

Suppose XX is some quantity measured on members of a population: the mass of a loaf, the score on a spinner, the lifetime of a bulb. A random sample of size nn gives nn observations X1,X2,…,XnX_1, X_2, \dots, X_n. Because the sample is random, these are independent and each has the same distribution as XX. The sample mean is

Xˉ=X1+X2+⋯+Xnn.\bar{X} = \frac{X_1 + X_2 + \cdots + X_n}{n}.

The capital letter matters. Xˉ\bar{X} is the random variable: the mean of a sample not yet taken. xˉ\bar{x} is a number: the mean of the sample you actually took.

A tiny example makes this concrete. A spinner gives 11, 22 or 33, each with probability 13\tfrac{1}{3}. Spin it twice and take the mean. The nine equally likely outcomes give:

Spins(1,1)(1,1)(1,2),(2,1)(1,2), (2,1)(1,3),(2,2),(3,1)(1,3), (2,2), (3,1)(2,3),(3,2)(2,3), (3,2)(3,3)(3,3)
xˉ\bar{x}111.51.5222.52.533
P(Xˉ=xˉ)P(\bar{X} = \bar{x})19\tfrac{1}{9}29\tfrac{2}{9}39\tfrac{3}{9}29\tfrac{2}{9}19\tfrac{1}{9}

A single spin has mean 22 and variance E(X2)−22=143−4=23E(X^2) - 2^2 = \tfrac{14}{3} - 4 = \tfrac{2}{3}. The sample mean also has mean 22, but its variance is

E(Xˉ2)−22=1+2(2.25)+3(4)+2(6.25)+99−4=399−4=13.E(\bar{X}^2) - 2^2 = \frac{1 + 2(2.25) + 3(4) + 2(6.25) + 9}{9} - 4 = \frac{39}{9} - 4 = \frac{1}{3}.

That is exactly half of 23\tfrac{2}{3}, and the sample size was 22. The extreme values 11 and 33 need both spins to be extreme, so they are rarer for the mean than for a single spin, and the distribution of Xˉ\bar{X} is pulled in towards the centre. This is the whole story in miniature.

Mean and variance of Xˉ\bar{X}

The general result follows from the rules for linear combinations of random variables. Let XX have mean μ\mu and variance σ2\sigma^2. Then

E(Xˉ)=1n(E(X1)+⋯+E(Xn))=1n(nμ)=μ,E(\bar{X}) = \frac{1}{n}\big(E(X_1) + \cdots + E(X_n)\big) = \frac{1}{n}(n\mu) = \mu,

and, because the XiX_i are independent,

Var(Xˉ)=1n2(Var(X1)+⋯+Var(Xn))=1n2(nσ2)=σ2n.\text{Var}(\bar{X}) = \frac{1}{n^2}\big(\text{Var}(X_1) + \cdots + \text{Var}(X_n)\big) = \frac{1}{n^2}(n\sigma^2) = \frac{\sigma^2}{n}.
Key result

For a random sample of size nn from a population with mean μ\mu and variance σ2\sigma^2:

E(Xˉ)=μ,Var(Xˉ)=σ2n,standard deviation of Xˉ=σn.E(\bar{X}) = \mu, \qquad \text{Var}(\bar{X}) = \frac{\sigma^2}{n}, \qquad \text{standard deviation of } \bar{X} = \frac{\sigma}{\sqrt{n}}.

For the sample total T=X1+X2+⋯+XnT = X_1 + X_2 + \cdots + X_n:

E(T)=nμ,Var(T)=nσ2.E(T) = n\mu, \qquad \text{Var}(T) = n\sigma^2.

Two things to notice. First, the sample mean is centred on the population mean: on average it gets the right answer, which is why xˉ\bar{x} is a sensible estimate of μ\mu. Second, the spread shrinks as nn grows, but only with n\sqrt{n}. To halve the standard deviation of the mean you need four times as many observations.

The standard deviation σn\dfrac{\sigma}{\sqrt{n}} is often called the standard error of the mean. The syllabus does not require the term, but you will meet it in textbooks.

Sum of a sample is not a multiple of one value

The total of nn independent observations has variance nσ2n\sigma^2, not n2σ2n^2\sigma^2. The expression nXnX means one observation multiplied by nn, with variance n2σ2n^2\sigma^2. Six eggs in a box are six different eggs, so their total is X1+⋯+X6X_1 + \cdots + X_6, not 6X6X.

When the population is normal

The mean and variance of Xˉ\bar{X} hold for any population. The shape of the distribution needs one more fact. A sum of independent normal variables is normal (see linear combinations of normal variables), and dividing by nn keeps it normal. So:

Key result

If X∼N(μ,σ2)X \sim N(\mu, \sigma^2), then for a random sample of any size nn

Xˉ∼N(μ,σ2n)exactly,T=∑Xi∼N(nμ,nσ2).\bar{X} \sim N\left(\mu, \frac{\sigma^2}{n}\right) \quad \text{exactly}, \qquad T = \sum X_i \sim N(n\mu, n\sigma^2).

The graph shows a population X∼N(80,102)X \sim N(80, 10^2) (the widest curve) with the distributions of the sample mean for n=4n = 4, where Xˉ∼N(80,52)\bar{X} \sim N(80, 5^2), and for n=16n = 16, where Xˉ∼N(80,2.52)\bar{X} \sim N(80, 2.5^2) (the tallest curve). All three are centred on 8080; the bigger the sample, the more tightly the mean clusters around 8080.

y = exp(-(x - 80)^2 / 200) / (10 sqrt(2 pi)) y = exp(-(x - 80)^2 / 50) / (5 sqrt(2 pi)) y = exp(-(x - 80)^2 / 12.5) / (2.5 sqrt(2 pi))

If the population is not normal, Xˉ\bar{X} still has mean μ\mu and variance σ2n\dfrac{\sigma^2}{n}, but its shape is not exactly normal. For a large sample it is approximately normal anyway: that is the Central Limit Theorem, the next note. For a small sample from a non-normal population, Paper 6 cannot ask you for probabilities about Xˉ\bar{X} beyond listing outcomes as in the spinner example.

Probabilities for a sample mean
  1. Define the variable and state the population distribution: "X∼N(800,152)X \sim N(800, 15^2), where XX is the mass of a loaf in grams".
  2. Write down the distribution of Xˉ\bar{X}, with the variance divided by nn: "Xˉ∼N(800,1529)\bar{X} \sim N\left(800, \tfrac{15^2}{9}\right)". Say why it is normal.
  3. Standardise with the standard deviation of the mean: z=xˉ−μσ/nz = \dfrac{\bar{x} - \mu}{\sigma/\sqrt{n}}.
  4. Use the normal tables and a sketch to find the probability.
Routine: one sample mean

The masses of loaves from a bakery are normally distributed with mean 800800 g and standard deviation 1515 g.

(a) Find the probability that a single randomly chosen loaf has mass less than 790790 g.

(b) Find the probability that the mean mass of a random sample of 99 loaves is less than 790790 g.

(c) Explain why the answer to (b) is smaller than the answer to (a).

Solution

Let XX g be the mass of a loaf, X∼N(800,152)X \sim N(800, 15^2).

(a)

P(X<790)=P(Z<790−80015)=P(Z<−0.667)=1−0.7476=0.252P(X < 790) = P\left(Z < \frac{790 - 800}{15}\right) = P(Z < -0.667) = 1 - 0.7476 = 0.252

(b) The population is normal, so Xˉ∼N(800,1529)=N(800,25)\bar{X} \sim N\left(800, \dfrac{15^2}{9}\right) = N(800, 25), with standard deviation 55.

P(Xˉ<790)=P(Z<790−8005)=P(Z<−2)=1−0.9772=0.0228P(\bar{X} < 790) = P\left(Z < \frac{790 - 800}{5}\right) = P(Z < -2) = 1 - 0.9772 = 0.0228

(c) The sample mean has a smaller variance than a single loaf. For the mean to be below 790790 g, the nine loaves would have to be light on average; one heavy loaf would pull the mean back up. So a low mean is much less likely than one light loaf.

A discrete population

The random variable XX has the following probability distribution.

xx00112233
P(X=x)P(X = x)0.10.10.30.30.40.40.20.2

(a) Find E(Xˉ)E(\bar{X}) and Var(Xˉ)\text{Var}(\bar{X}), where Xˉ\bar{X} is the mean of a random sample of 1010 observations of XX.

(b) A random sample of 22 observations is taken. Find P(Xˉ=1)P(\bar{X} = 1) and P(Xˉ>2.5)P(\bar{X} > 2.5).

Solution

(a)

E(X)=0(0.1)+1(0.3)+2(0.4)+3(0.2)=1.7E(X) = 0(0.1) + 1(0.3) + 2(0.4) + 3(0.2) = 1.7E(X2)=0+0.3+1.6+1.8=3.7,Var(X)=3.7−1.72=0.81E(X^2) = 0 + 0.3 + 1.6 + 1.8 = 3.7, \qquad \text{Var}(X) = 3.7 - 1.7^2 = 0.81E(Xˉ)=1.7,Var(Xˉ)=0.8110=0.081E(\bar{X}) = 1.7, \qquad \text{Var}(\bar{X}) = \frac{0.81}{10} = 0.081

(b) Xˉ=1\bar{X} = 1 means the two values add to 22: (0,2)(0, 2), (2,0)(2, 0) or (1,1)(1, 1).

P(Xˉ=1)=0.1(0.4)+0.4(0.1)+0.3(0.3)=0.04+0.04+0.09=0.17P(\bar{X} = 1) = 0.1(0.4) + 0.4(0.1) + 0.3(0.3) = 0.04 + 0.04 + 0.09 = 0.17

Xˉ>2.5\bar{X} > 2.5 needs a total greater than 55, which is only (3,3)(3, 3).

P(Xˉ>2.5)=0.22=0.04P(\bar{X} > 2.5) = 0.2^2 = 0.04

Note that the population is not normal and the sample in (b) is tiny, so the only way to find these probabilities is to list the outcomes.

Finding a sample size

The time taken by a machine to complete a cycle is normally distributed with mean 5050 seconds and standard deviation 66 seconds. Find the smallest sample size nn for which the probability that the mean time of the sample exceeds 5252 seconds is less than 0.010.01.

Solution

Xˉ∼N(50,36n)\bar{X} \sim N\left(50, \dfrac{36}{n}\right), so the standard deviation of Xˉ\bar{X} is 6n\dfrac{6}{\sqrt{n}}.

We need

P(Xˉ>52)<0.01⟺52−506/n>2.326.P(\bar{X} > 52) < 0.01 \quad\Longleftrightarrow\quad \frac{52 - 50}{6/\sqrt{n}} > 2.326.

The value 2.3262.326 comes from the table of critical values: Φ(2.326)=0.99\Phi(2.326) = 0.99.

2n6>2.326  ⇒  n>6.978  ⇒  n>48.69\frac{2\sqrt{n}}{6} > 2.326 \;\Rightarrow\; \sqrt{n} > 6.978 \;\Rightarrow\; n > 48.69

The smallest sample size is n=49n = 49.

Check the direction: a larger nn makes the mean less spread out, so the probability of it straying above 5252 falls. The inequality n>48.69n > 48.69 is consistent with that.

Working backwards to the mean

Rods are cut to lengths that are normally distributed with mean μ\mu cm and standard deviation 0.40.4 cm. For a random sample of 2525 rods, the probability that the sample mean exceeds 20.120.1 cm is 0.050.05. Find μ\mu.

Solution

Xˉ∼N(μ,0.4225)\bar{X} \sim N\left(\mu, \dfrac{0.4^2}{25}\right), standard deviation 0.45=0.08\dfrac{0.4}{5} = 0.08.

P(Xˉ>20.1)=0.05P(\bar{X} > 20.1) = 0.05, so 20.120.1 is 1.6451.645 standard deviations above the mean:

20.1−μ0.08=1.645  ⇒  μ=20.1−0.1316=19.97 (4 s.f.)\frac{20.1 - \mu}{0.08} = 1.645 \;\Rightarrow\; \mu = 20.1 - 0.1316 = 19.97 \text{ (4 s.f.)}
Exam-hard: comparing two samples

The masses of apples of variety A are normally distributed with mean 150150 g and standard deviation 2020 g. The masses of apples of variety B are normally distributed with mean 140140 g and standard deviation 1515 g. A random sample of 88 apples of variety A and a random sample of 1010 apples of variety B are chosen.

(a) Find the probability that the mean mass of the A apples is greater than the mean mass of the B apples.

(b) Find the probability that the total mass of the 88 A apples is greater than the total mass of the 1010 B apples.

Solution

(a)

Aˉ∼N(150,4008)=N(150,50),Bˉ∼N(140,22510)=N(140,22.5)\bar{A} \sim N\left(150, \frac{400}{8}\right) = N(150, 50), \qquad \bar{B} \sim N\left(140, \frac{225}{10}\right) = N(140, 22.5)

The samples are independent, so

Aˉ−Bˉ∼N(150−140, 50+22.5)=N(10,72.5).\bar{A} - \bar{B} \sim N(150 - 140,\ 50 + 22.5) = N(10, 72.5).P(Aˉ−Bˉ>0)=P(Z>0−1072.5)=P(Z>−1.174)=0.880P(\bar{A} - \bar{B} > 0) = P\left(Z > \frac{0 - 10}{\sqrt{72.5}}\right) = P(Z > -1.174) = 0.880

(b) Totals: TA∼N(8×150, 8×400)=N(1200,3200)T_A \sim N(8 \times 150,\ 8 \times 400) = N(1200, 3200) and TB∼N(10×140, 10×225)=N(1400,2250)T_B \sim N(10 \times 140,\ 10 \times 225) = N(1400, 2250).

TA−TB∼N(−200, 5450)T_A - T_B \sim N(-200,\ 5450)P(TA−TB>0)=P(Z>2005450)=P(Z>2.709)=1−0.99663=0.00337P(T_A - T_B > 0) = P\left(Z > \frac{200}{\sqrt{5450}}\right) = P(Z > 2.709) = 1 - 0.99663 = 0.00337

The variances add in both parts even though the variables are subtracted. Part (b) is very different from part (a) because there are more B apples, so the B total is larger on average.

Common mistakes
  • Using σ\sigma instead of σn\dfrac{\sigma}{\sqrt{n}}. When the question is about a sample mean, standardise with the standard deviation of the mean. This is the single most common error in the topic.
  • Dividing the standard deviation by nn. The variance is divided by nn; the standard deviation is divided by n\sqrt{n}. N(800,1529)N\left(800, \tfrac{15^2}{9}\right) has standard deviation 55, not 159\tfrac{15}{9}.
  • Writing N(800,15)N(800, 15) when the standard deviation is 1515. The second parameter of N(μ,σ2)N(\mu, \sigma^2) is the variance. Write N(800,152)N(800, 15^2) or N(800,225)N(800, 225).
  • Treating a total as nXnX. The total of nn observations has variance nσ2n\sigma^2, not n2σ2n^2\sigma^2.
  • Claiming Xˉ\bar{X} is normal for a small sample from a non-normal population. It is exactly normal only if XX is normal; otherwise you need a large sample and the Central Limit Theorem.
  • Subtracting variances. For Aˉ−Bˉ\bar{A} - \bar{B} the variances add.
Exam tip
  • Always write the distribution of Xˉ\bar{X} in full before standardising, for example "Xˉ∼N(800,2259)\bar{X} \sim N\left(800, \tfrac{225}{9}\right)". This line usually carries a mark of its own.
  • When the population is stated to be normal, say "since XX is normal, Xˉ\bar{X} is normal". Do not quote the Central Limit Theorem here: the result is exact, and examiners penalise quoting the CLT when it is not needed.
  • "Find the smallest nn" questions: set up the inequality, solve it, and round up to the next integer, whatever the decimal part.
  • Use the critical values printed under the normal table (1.6451.645, 1.961.96, 2.3262.326, 2.5762.576) rather than reading them inexactly from the main table.
  • Give probabilities to 3 significant figures, keeping 4 or more figures in intermediate values such as zz.
Summary
  • Xˉ\bar{X}, the mean of a random sample of size nn, is a random variable.
  • E(Xˉ)=μE(\bar{X}) = \mu and Var(Xˉ)=σ2n\text{Var}(\bar{X}) = \dfrac{\sigma^2}{n} for any population; the standard deviation of Xˉ\bar{X} is σn\dfrac{\sigma}{\sqrt{n}}.
  • The sample total has mean nμn\mu and variance nσ2n\sigma^2.
  • If XX is normal, Xˉ∼N(μ,σ2n)\bar{X} \sim N\left(\mu, \dfrac{\sigma^2}{n}\right) exactly, for every nn.
  • Standardise a sample mean with z=xˉ−μσ/nz = \dfrac{\bar{x} - \mu}{\sigma/\sqrt{n}}.
  • Larger samples give means that cluster more tightly around μ\mu; quadrupling nn halves the spread.
  • For differences of sample means or totals, the means subtract and the variances add.

Practice questions

Question
  1. X∼N(24,52)X \sim N(24, 5^2). A random sample of 1010 observations of XX is taken. State the distribution of Xˉ\bar{X} and find P(Xˉ>25.5)P(\bar{X} > 25.5).
  2. The random variable XX takes the values 11, 22 and 55 with probabilities 0.50.5, 0.30.3 and 0.20.2 respectively. Find the mean and variance of the mean of a random sample of 4040 observations of XX.
  3. The masses of eggs are normally distributed with mean 6262 g and standard deviation 44 g. Eggs are packed in boxes of 66. Find the probability that the total mass of eggs in a randomly chosen box is less than 360360 g.
  4. X∼N(μ,32)X \sim N(\mu, 3^2). For a random sample of 1616 observations, P(Xˉ<48.2)=0.1P(\bar{X} < 48.2) = 0.1. Find μ\mu.
  5. IQ scores are normally distributed with mean 100100 and standard deviation 1515. Find the smallest sample size for which the probability that the sample mean lies within 33 of 100100 is at least 0.950.95.
  6. The heights of a species of plant are normally distributed with mean 7070 cm and standard deviation 1212 cm. Find the probability that the mean height of a random sample of 99 plants lies between 6666 cm and 7575 cm.
  7. X∼N(30,22)X \sim N(30, 2^2) and Y∼N(29,32)Y \sim N(29, 3^2) are independent. A random sample of 55 observations of XX and a random sample of 44 observations of YY are taken. Find the probability that Xˉ>Yˉ\bar{X} > \bar{Y}.
  8. X∼N(20,16)X \sim N(20, 16). A random sample of nn observations of XX is taken, and P(Xˉ>21)=0.0668P(\bar{X} > 21) = 0.0668. (a) Find nn. (b) Find the probability that the sum of the nn observations exceeds 740740.
Answers
  1. Since XX is normal, Xˉ∼N(24,2510)=N(24,2.5)\bar{X} \sim N\left(24, \dfrac{25}{10}\right) = N(24, 2.5). P(Xˉ>25.5)=P(Z>1.52.5)=P(Z>0.949)=1−0.8287=0.171P(\bar{X} > 25.5) = P\left(Z > \dfrac{1.5}{\sqrt{2.5}}\right) = P(Z > 0.949) = 1 - 0.8287 = 0.171.

  2. E(X)=0.5+0.6+1.0=2.1E(X) = 0.5 + 0.6 + 1.0 = 2.1; E(X2)=0.5+1.2+5=6.7E(X^2) = 0.5 + 1.2 + 5 = 6.7; Var(X)=6.7−2.12=2.29\text{Var}(X) = 6.7 - 2.1^2 = 2.29. E(Xˉ)=2.1E(\bar{X}) = 2.1 and Var(Xˉ)=2.2940=0.0573\text{Var}(\bar{X}) = \dfrac{2.29}{40} = 0.0573 (3 s.f.).

  3. T=X1+⋯+X6∼N(6×62, 6×16)=N(372,96)T = X_1 + \cdots + X_6 \sim N(6 \times 62,\ 6 \times 16) = N(372, 96). P(T<360)=P(Z<360−37296)=P(Z<−1.225)=1−0.8897=0.110P(T < 360) = P\left(Z < \dfrac{360 - 372}{\sqrt{96}}\right) = P(Z < -1.225) = 1 - 0.8897 = 0.110.

  4. Xˉ∼N(μ,916)\bar{X} \sim N\left(\mu, \dfrac{9}{16}\right), standard deviation 0.750.75. Since P(Xˉ<48.2)=0.1P(\bar{X} < 48.2) = 0.1, 48.248.2 is below the mean: 48.2−μ0.75=−1.282\dfrac{48.2 - \mu}{0.75} = -1.282, so μ=48.2+0.9615=49.16\mu = 48.2 + 0.9615 = 49.16 (4 s.f.).

  5. Xˉ∼N(100,225n)\bar{X} \sim N\left(100, \dfrac{225}{n}\right). P(97<Xˉ<103)≥0.95P(97 < \bar{X} < 103) \ge 0.95 needs 315/n≥1.96\dfrac{3}{15/\sqrt{n}} \ge 1.96, so n≥9.8\sqrt{n} \ge 9.8 and n≥96.04n \ge 96.04. The smallest sample size is 9797.

  6. Xˉ∼N(70,1449)=N(70,16)\bar{X} \sim N\left(70, \dfrac{144}{9}\right) = N(70, 16), standard deviation 44. P(66<Xˉ<75)=P(−1<Z<1.25)=0.8944−(1−0.8413)=0.8944−0.1587=0.736P(66 < \bar{X} < 75) = P(-1 < Z < 1.25) = 0.8944 - (1 - 0.8413) = 0.8944 - 0.1587 = 0.736.

  7. Xˉ∼N(30,45)\bar{X} \sim N\left(30, \dfrac{4}{5}\right) and Yˉ∼N(29,94)\bar{Y} \sim N\left(29, \dfrac{9}{4}\right), so Xˉ−Yˉ∼N(1, 0.8+2.25)=N(1,3.05)\bar{X} - \bar{Y} \sim N(1,\ 0.8 + 2.25) = N(1, 3.05). P(Xˉ−Yˉ>0)=P(Z>−13.05)=P(Z>−0.5726)=0.717P(\bar{X} - \bar{Y} > 0) = P\left(Z > \dfrac{-1}{\sqrt{3.05}}\right) = P(Z > -0.5726) = 0.717.

  8. (a) Xˉ∼N(20,16n)\bar{X} \sim N\left(20, \dfrac{16}{n}\right). P(Xˉ>21)=0.0668P(\bar{X} > 21) = 0.0668 means Φ(z)=0.9332\Phi(z) = 0.9332, so z=1.5z = 1.5. 21−204/n=1.5⇒n=6⇒n=36\dfrac{21 - 20}{4/\sqrt{n}} = 1.5 \Rightarrow \sqrt{n} = 6 \Rightarrow n = 36. (b) T∼N(36×20, 36×16)=N(720,576)T \sim N(36 \times 20,\ 36 \times 16) = N(720, 576), standard deviation 2424. P(T>740)=P(Z>2024)=P(Z>0.833)=1−0.7976=0.202P(T > 740) = P\left(Z > \dfrac{20}{24}\right) = P(Z > 0.833) = 1 - 0.7976 = 0.202. (Equivalently, T>740T > 740 is the same event as Xˉ>74036=20.56\bar{X} > \tfrac{740}{36} = 20.56.)

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action