Unbiased Estimates of Mean and Variance

A2 · S2 · 14 min

The population mean μ\mu and variance σ2\sigma^2 are almost never known. What you have is a sample, and from it you must produce the best single-number guesses for μ\mu and σ2\sigma^2. The sample mean does the job for μ\mu. For σ2\sigma^2 there is a catch: the variance you learnt in S1, which divides by nn, comes out too small on average, so Paper 6 uses a version that divides by n−1n - 1. Almost every confidence interval and test question starts with this calculation, often from summarised or coded totals, and it is worth several easy marks if done carefully.

Estimators and estimates

A parameter such as μ\mu is a fixed but unknown number describing the population. To estimate it you use a formula applied to sample data. Before the sample is taken, that formula gives a random variable called an estimator; once the data are in, it gives a number called an estimate.

For example, the estimator of μ\mu is the sample mean Xˉ\bar{X}, a random variable. If your sample of loaves has mean 803.2803.2 g, then 803.2803.2 is the estimate.

Because an estimator is a random variable, it has a distribution. Some samples give estimates above the true value, some below. The question is whether the method is right on average.

Definition

An estimator is unbiased if its expected value equals the population parameter it estimates. Although individual estimates vary from sample to sample, the process gives the correct value on average.

An estimate calculated using an unbiased estimator is called an unbiased estimate.

The syllabus only asks for this simple understanding. You will not be asked to prove that an estimator is unbiased.

Estimating the mean

You know from the distribution of the sample mean that E(Xˉ)=μE(\bar{X}) = \mu. So the sample mean is an unbiased estimator of the population mean, and

xˉ=∑xn\bar{x} = \frac{\sum x}{n}

is the unbiased estimate of μ\mu. Nothing new is needed.

Estimating the variance: why n−1n - 1

The natural guess for σ2\sigma^2 is the variance of the sample, calculated as in S1:

∑(x−xˉ)2n.\frac{\sum (x - \bar{x})^2}{n}.

This is biased. It measures spread around xˉ\bar{x}, the centre of the sample, and xˉ\bar{x} is always exactly in the middle of the sample values. The true mean μ\mu is usually somewhere else, and the sum of squared deviations is smallest when measured from xˉ\bar{x}. So deviations from xˉ\bar{x} are, on average, a little too small, and the estimate of the variance is too small.

A small example shows the size of the effect. A spinner gives 11, 22 or 33 with equal probability, so its variance is σ2=23\sigma^2 = \tfrac{2}{3}. Take samples of size n=2n = 2. For a sample (a,b)(a, b) the variance dividing by nn is (a−b2)2\left(\dfrac{a - b}{2}\right)^2, and dividing by n−1n - 1 it is (a−b)22\dfrac{(a - b)^2}{2}. Over the nine equally likely samples:

Difference ∣a−b∣\lvert a - b\rvert001122
Number of samples334422
Variance, dividing by nn000.250.2511
Variance, dividing by n−1n - 1000.50.522

The averages are

3(0)+4(0.25)+2(1)9=13and3(0)+4(0.5)+2(2)9=23.\frac{3(0) + 4(0.25) + 2(1)}{9} = \frac{1}{3} \qquad\text{and}\qquad \frac{3(0) + 4(0.5) + 2(2)}{9} = \frac{2}{3}.

Dividing by nn gives 13\tfrac{1}{3} on average, which is too small. Dividing by n−1n - 1 gives 23\tfrac{2}{3} on average: exactly σ2\sigma^2. In general, dividing by nn gives on average n−1nσ2\dfrac{n - 1}{n}\sigma^2, and multiplying by nn−1\dfrac{n}{n - 1} removes the bias.

Unbiased estimates

From a random sample x1,x2,…,xnx_1, x_2, \dots, x_n:

xˉ=∑xn\bar{x} = \frac{\sum x}{n}s2=1n−1∑(x−xˉ)2=1n−1(∑x2−(∑x)2n)s^2 = \frac{1}{n - 1}\sum (x - \bar{x})^2 = \frac{1}{n - 1}\left(\sum x^2 - \frac{\left(\sum x\right)^2}{n}\right)

xˉ\bar{x} is the unbiased estimate of μ\mu and s2s^2 is the unbiased estimate of σ2\sigma^2. Both formulae are in the formula booklet.

If the variance of the sample (dividing by nn) is already known, then

s2=nn−1×(variance of the sample).s^2 = \frac{n}{n - 1} \times (\text{variance of the sample}).

For large samples the factor nn−1\dfrac{n}{n - 1} is close to 11 and the two versions are nearly equal. For small samples the difference is significant, and in every case the mark scheme expects n−1n - 1.

The forms the data come in

Questions give the data in one of four ways. The method is always the same: get nn, ∑x\sum x and ∑x2\sum x^2 (or their coded versions), then substitute.

Raw data. Add up the values and their squares, or use the statistics mode of your calculator. Most calculators show two standard deviations: one labelled σn\sigma_n or σx\sigma_x (dividing by nn) and one labelled sn−1s_{n-1} or sxs_x (dividing by n−1n - 1). Use the second, and square it.

Summary totals. You are given nn, ∑x\sum x and ∑x2\sum x^2. Substitute directly.

Coded totals. You are given ∑(x−a)\sum (x - a) and ∑(x−a)2\sum (x - a)^2 for some constant aa. Subtracting a constant shifts every value by the same amount, which moves the mean but leaves the spread unchanged. So

xˉ=a+∑(x−a)n,s2=1n−1(∑(x−a)2−(∑(x−a))2n).\bar{x} = a + \frac{\sum (x - a)}{n}, \qquad s^2 = \frac{1}{n - 1}\left(\sum (x - a)^2 - \frac{\left(\sum (x - a)\right)^2}{n}\right).

A frequency table. Use ∑fx\sum fx and ∑fx2\sum fx^2 in place of ∑x\sum x and ∑x2\sum x^2, with n=∑fn = \sum f.

Unbiased estimates of mean and variance
  1. Write down nn, ∑x\sum x and ∑x2\sum x^2 (or ∑(x−a)\sum (x - a) and ∑(x−a)2\sum (x - a)^2, or ∑fx\sum fx and ∑fx2\sum fx^2).
  2. Mean: xˉ=∑xn\bar{x} = \dfrac{\sum x}{n}, adding back aa if the data are coded.
  3. Variance: s2=1n−1(∑x2−(∑x)2n)s^2 = \dfrac{1}{n - 1}\left(\sum x^2 - \dfrac{(\sum x)^2}{n}\right), with coded totals in place of ∑x\sum x and ∑x2\sum x^2 if given. Never add aa to the variance.
  4. Give answers to 3 significant figures (or exactly), and keep more figures if the value is used later.
Raw data

A random sample of 66 packets of rice has masses, in grams,

498,503,501,497,505,502.498,\quad 503,\quad 501,\quad 497,\quad 505,\quad 502.

Find unbiased estimates of the population mean and variance.

Solution

n=6n = 6, ∑x=3006\sum x = 3006, ∑x2=1 506 052\sum x^2 = 1\,506\,052.

xˉ=30066=501\bar{x} = \frac{3006}{6} = 501s2=15(1 506 052−300626)=15(1 506 052−1 506 006)=465=9.2s^2 = \frac{1}{5}\left(1\,506\,052 - \frac{3006^2}{6}\right) = \frac{1}{5}(1\,506\,052 - 1\,506\,006) = \frac{46}{5} = 9.2

Check with deviations from 501501: −3,2,0,−4,4,1-3, 2, 0, -4, 4, 1, whose squares add to 9+4+0+16+16+1=469 + 4 + 0 + 16 + 16 + 1 = 46. Dividing by 55 gives 9.29.2. With large numbers like these, the deviation form avoids subtracting two huge, nearly equal totals.

Summary totals and coded totals

(a) For a random sample of 5050 values of xx, ∑x=1240\sum x = 1240 and ∑x2=31 900\sum x^2 = 31\,900. Find unbiased estimates of the population mean and variance.

(b) The times, tt seconds, taken by a random sample of 4040 students to solve a puzzle are summarised by ∑(t−30)=64\sum (t - 30) = 64 and ∑(t−30)2=1130\sum (t - 30)^2 = 1130. Find unbiased estimates of the population mean and variance of tt.

Solution

(a)

xˉ=124050=24.8,s2=149(31 900−1240250)=149(31 900−30 752)=114849=23.4\bar{x} = \frac{1240}{50} = 24.8, \qquad s^2 = \frac{1}{49}\left(31\,900 - \frac{1240^2}{50}\right) = \frac{1}{49}(31\,900 - 30\,752) = \frac{1148}{49} = 23.4

(b) The mean of the coded values is 6440=1.6\dfrac{64}{40} = 1.6, so

tˉ=30+1.6=31.6.\bar{t} = 30 + 1.6 = 31.6.

The variance is unaffected by subtracting 3030:

s2=139(1130−64240)=139(1130−102.4)=1027.639=26.3s^2 = \frac{1}{39}\left(1130 - \frac{64^2}{40}\right) = \frac{1}{39}(1130 - 102.4) = \frac{1027.6}{39} = 26.3
A frequency table

The numbers of goals scored in a random sample of 5050 football matches are shown.

Goals, xx001122334455
Frequency, ff8815151212994422

(a) Find unbiased estimates of the mean and variance of the number of goals per match.

(b) Comment on whether a Poisson distribution could be a suitable model.

Solution

(a)

∑fx=0+15+24+27+16+10=92,∑fx2=0+15+48+81+64+50=258\sum fx = 0 + 15 + 24 + 27 + 16 + 10 = 92, \qquad \sum fx^2 = 0 + 15 + 48 + 81 + 64 + 50 = 258xˉ=9250=1.84,s2=149(258−92250)=149(258−169.28)=88.7249=1.81\bar{x} = \frac{92}{50} = 1.84, \qquad s^2 = \frac{1}{49}\left(258 - \frac{92^2}{50}\right) = \frac{1}{49}(258 - 169.28) = \frac{88.72}{49} = 1.81

(b) For a Poisson distribution the mean equals the variance. The estimates 1.841.84 and 1.811.81 are close, so a Poisson model is plausible (see modelling with the Poisson distribution).

From a calculator's standard deviation

A calculator gives the standard deviation of a sample of 2020 values, calculated by dividing by nn, as 4.54.5. Find the unbiased estimate of the population variance.

Solution

The variance of the sample is 4.52=20.254.5^2 = 20.25, so

s2=2019×20.25=21.3 (3 s.f.)s^2 = \frac{20}{19} \times 20.25 = 21.3 \text{ (3 s.f.)}
Exam-hard: combining two samples

A random sample of 2020 observations of a random variable XX has mean 14.214.2, and the unbiased estimate of the population variance from this sample is 3.63.6. A second, independent random sample of 3030 observations of XX gives ∑x=441\sum x = 441 and ∑x2=6600\sum x^2 = 6600.

(a) Using both samples together, find unbiased estimates of the mean and variance of XX.

(b) Use your estimates to find the approximate probability that the mean of a further random sample of 4040 observations of XX is greater than 1515.

Solution

(a) Recover the totals for the first sample.

∑x=20×14.2=284\sum x = 20 \times 14.2 = 284

From s2=1n−1(∑x2−(∑x)2n)s^2 = \dfrac{1}{n - 1}\left(\sum x^2 - \dfrac{(\sum x)^2}{n}\right):

3.6=119(∑x2−284220)  ⇒  ∑x2=19×3.6+4032.8=4101.23.6 = \frac{1}{19}\left(\sum x^2 - \frac{284^2}{20}\right) \;\Rightarrow\; \sum x^2 = 19 \times 3.6 + 4032.8 = 4101.2

Combined: n=50n = 50, ∑x=284+441=725\sum x = 284 + 441 = 725, ∑x2=4101.2+6600=10 701.2\sum x^2 = 4101.2 + 6600 = 10\,701.2.

xˉ=72550=14.5\bar{x} = \frac{725}{50} = 14.5s2=149(10 701.2−725250)=149(10 701.2−10 512.5)=188.749=3.851s^2 = \frac{1}{49}\left(10\,701.2 - \frac{725^2}{50}\right) = \frac{1}{49}(10\,701.2 - 10\,512.5) = \frac{188.7}{49} = 3.851

(b) The distribution of XX is unknown, but n=40n = 40 is large, so by the Central Limit Theorem

Xˉ∼N(14.5,3.85140) approximately,standard deviation 0.31028.\bar{X} \sim N\left(14.5, \frac{3.851}{40}\right) \text{ approximately}, \qquad \text{standard deviation } 0.31028.P(Xˉ>15)=P(Z>0.50.31028)=P(Z>1.611)=1−0.9464=0.0536P(\bar{X} > 15) = P\left(Z > \frac{0.5}{0.31028}\right) = P(Z > 1.611) = 1 - 0.9464 = 0.0536

You cannot simply average the two variance estimates: the samples have different sizes and different means, and the spread of the combined data includes the gap between those means.

Where the estimates are used

When σ2\sigma^2 is unknown and the sample is large, s2s^2 stands in for σ2\sigma^2 in confidence intervals and in tests for a population mean. An error in s2s^2 therefore carries through the rest of a long question, which is why the method mark for this step is so valuable.

In the same spirit, the sample proportion ps=xnp_s = \dfrac{x}{n} is an unbiased estimate of a population proportion pp, because E(Xn)=npn=pE\left(\dfrac{X}{n}\right) = \dfrac{np}{n} = p when X∼B(n,p)X \sim B(n, p). It is used in confidence intervals for proportions.

Common mistakes
  • Dividing by nn. The unbiased estimate of variance divides by n−1n - 1. Using nn loses the accuracy mark, and everything built on it is wrong.
  • Using the wrong calculator value. σn\sigma_n (or σx\sigma_x) divides by nn; sn−1s_{n-1} (or sxs_x) divides by n−1n - 1. Square the right one.
  • Adding the coding constant to the variance. ∑(x−a)\sum (x - a) needs aa added back to get the mean; the variance needs nothing added.
  • Mixing coded and uncoded totals. In ∑(x−a)2−(∑(x−a))2n\sum (x - a)^2 - \dfrac{(\sum (x - a))^2}{n} both totals must be coded.
  • Forgetting to square ∑x\sum x. It is (∑x)2n\dfrac{(\sum x)^2}{n}, not ∑x2n\dfrac{\sum x^2}{n}.
  • Giving ss when asked for the variance, or s2s^2 when a standard deviation is needed in the next part. Read which one is asked for.
  • Explaining "unbiased" as "the sample is random" or "the estimate is accurate". It means the expected value of the estimator equals the parameter, so the method is right on average.
Exam tip
  • The question words "unbiased estimates of the population mean and variance" signal exactly this calculation. Present it as two separate lines, formula then numbers.
  • Show the substitution into the formula, such as 149(31 900−1240250)\dfrac{1}{49}\left(31\,900 - \dfrac{1240^2}{50}\right). A slip after a correct substitution keeps the method mark.
  • Give answers exactly where they terminate (9.29.2, 31.631.6) or to 3 significant figures. If the value is used later, carry more figures.
  • When a later part needs the standard deviation, take s=s2s = \sqrt{s^2} from the unrounded value.
  • "Explain what is meant by an unbiased estimate" is worth one mark: "the expected value of the estimator equals the population parameter, so on average it gives the true value".
Summary
  • An estimator is unbiased if its expected value equals the parameter: right on average, though individual estimates vary.
  • xˉ=∑xn\bar{x} = \dfrac{\sum x}{n} is the unbiased estimate of μ\mu.
  • s2=1n−1(∑x2−(∑x)2n)s^2 = \dfrac{1}{n - 1}\left(\sum x^2 - \dfrac{(\sum x)^2}{n}\right) is the unbiased estimate of σ2\sigma^2.
  • Dividing by nn underestimates σ2\sigma^2 on average, by the factor n−1n\dfrac{n - 1}{n}.
  • s2=nn−1×s^2 = \dfrac{n}{n - 1} \times (variance of the sample, dividing by nn).
  • For coded data, add aa back to the mean only; the variance is unchanged by coding.
  • For frequency tables use ∑fx\sum fx, ∑fx2\sum fx^2 and n=∑fn = \sum f.

Practice questions

Question
  1. A random sample gives the values 12,15,11,18,1412, 15, 11, 18, 14. Find unbiased estimates of the population mean and variance.

  2. For a random sample of 8080 values, ∑x=2000\sum x = 2000 and ∑x2=52 600\sum x^2 = 52\,600. Find unbiased estimates of the population mean and variance.

  3. For a random sample of 2525 values, ∑(x−100)=−30\sum (x - 100) = -30 and ∑(x−100)2=486\sum (x - 100)^2 = 486. Find unbiased estimates of the population mean and variance.

  4. The variance of a random sample of 1010 values, calculated by dividing by nn, is 6.36.3. Find the unbiased estimate of the population variance.

  5. The numbers of children in a random sample of 7070 households are shown.

    Children, xx0011223344
    Frequency121220202525101033

    Find unbiased estimates of the population mean and variance.

  6. A researcher says that the sample mean is an unbiased estimator of the population mean. Explain what this means.

  7. A random sample of 1212 observations has mean 5.55.5, and the unbiased estimate of the population variance is 2.42.4. A thirteenth observation, 88, is added to the sample. Find the new unbiased estimates of the population mean and variance.

  8. For a random sample of 4040 values of xx, ∑(x−k)=48\sum (x - k) = 48 and ∑(x−k)2=900\sum (x - k)^2 = 900, where kk is a constant. (a) Find the unbiased estimate of the population variance. (b) The unbiased estimate of the population mean is 21.221.2. Find kk.

Answers
  1. ∑x=70\sum x = 70, ∑x2=1010\sum x^2 = 1010, n=5n = 5. xˉ=14\bar{x} = 14; s2=14(1010−7025)=14(1010−980)=7.5s^2 = \tfrac{1}{4}\left(1010 - \tfrac{70^2}{5}\right) = \tfrac{1}{4}(1010 - 980) = 7.5.

  2. xˉ=200080=25\bar{x} = \tfrac{2000}{80} = 25; s2=179(52 600−2000280)=179(52 600−50 000)=260079=32.9s^2 = \tfrac{1}{79}\left(52\,600 - \tfrac{2000^2}{80}\right) = \tfrac{1}{79}(52\,600 - 50\,000) = \tfrac{2600}{79} = 32.9.

  3. xˉ=100+−3025=98.8\bar{x} = 100 + \tfrac{-30}{25} = 98.8; s2=124(486−(−30)225)=124(486−36)=18.75s^2 = \tfrac{1}{24}\left(486 - \tfrac{(-30)^2}{25}\right) = \tfrac{1}{24}(486 - 36) = 18.75.

  4. s2=109×6.3=7s^2 = \tfrac{10}{9} \times 6.3 = 7.

  5. ∑f=70\sum f = 70, ∑fx=0+20+50+30+12=112\sum fx = 0 + 20 + 50 + 30 + 12 = 112, ∑fx2=0+20+100+90+48=258\sum fx^2 = 0 + 20 + 100 + 90 + 48 = 258. xˉ=11270=1.6\bar{x} = \tfrac{112}{70} = 1.6; s2=169(258−112270)=169(258−179.2)=78.869=1.14s^2 = \tfrac{1}{69}\left(258 - \tfrac{112^2}{70}\right) = \tfrac{1}{69}(258 - 179.2) = \tfrac{78.8}{69} = 1.14.

  6. The expected value of the sample mean equals the population mean, E(Xˉ)=μE(\bar{X}) = \mu. Different samples give different sample means, some too high and some too low, but on average the sample mean equals the population mean.

  7. Original: ∑x=12×5.5=66\sum x = 12 \times 5.5 = 66 and ∑x2=11×2.4+66212=26.4+363=389.4\sum x^2 = 11 \times 2.4 + \tfrac{66^2}{12} = 26.4 + 363 = 389.4. New: n=13n = 13, ∑x=74\sum x = 74, ∑x2=389.4+64=453.4\sum x^2 = 389.4 + 64 = 453.4. xˉ=7413=5.69\bar{x} = \tfrac{74}{13} = 5.69; s2=112(453.4−74213)=112(453.4−421.23)=2.68s^2 = \tfrac{1}{12}\left(453.4 - \tfrac{74^2}{13}\right) = \tfrac{1}{12}(453.4 - 421.23) = 2.68.

  8. (a) s2=139(900−48240)=139(900−57.6)=842.439=21.6s^2 = \tfrac{1}{39}\left(900 - \tfrac{48^2}{40}\right) = \tfrac{1}{39}(900 - 57.6) = \tfrac{842.4}{39} = 21.6. Note that kk is not needed: coding does not affect the variance. (b) xˉ=k+4840=k+1.2=21.2\bar{x} = k + \tfrac{48}{40} = k + 1.2 = 21.2, so k=20k = 20.

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action