Unbiased Estimates of Mean and Variance
The population mean and variance are almost never known. What you have is a sample, and from it you must produce the best single-number guesses for and . The sample mean does the job for . For there is a catch: the variance you learnt in S1, which divides by , comes out too small on average, so Paper 6 uses a version that divides by . Almost every confidence interval and test question starts with this calculation, often from summarised or coded totals, and it is worth several easy marks if done carefully.
Estimators and estimates
A parameter such as is a fixed but unknown number describing the population. To estimate it you use a formula applied to sample data. Before the sample is taken, that formula gives a random variable called an estimator; once the data are in, it gives a number called an estimate.
For example, the estimator of is the sample mean , a random variable. If your sample of loaves has mean g, then is the estimate.
Because an estimator is a random variable, it has a distribution. Some samples give estimates above the true value, some below. The question is whether the method is right on average.
An estimator is unbiased if its expected value equals the population parameter it estimates. Although individual estimates vary from sample to sample, the process gives the correct value on average.
An estimate calculated using an unbiased estimator is called an unbiased estimate.
The syllabus only asks for this simple understanding. You will not be asked to prove that an estimator is unbiased.
Estimating the mean
You know from the distribution of the sample mean that . So the sample mean is an unbiased estimator of the population mean, and
is the unbiased estimate of . Nothing new is needed.
Estimating the variance: why
The natural guess for is the variance of the sample, calculated as in S1:
This is biased. It measures spread around , the centre of the sample, and is always exactly in the middle of the sample values. The true mean is usually somewhere else, and the sum of squared deviations is smallest when measured from . So deviations from are, on average, a little too small, and the estimate of the variance is too small.
A small example shows the size of the effect. A spinner gives , or with equal probability, so its variance is . Take samples of size . For a sample the variance dividing by is , and dividing by it is . Over the nine equally likely samples:
| Difference | |||
|---|---|---|---|
| Number of samples | |||
| Variance, dividing by | |||
| Variance, dividing by |
The averages are
Dividing by gives on average, which is too small. Dividing by gives on average: exactly . In general, dividing by gives on average , and multiplying by removes the bias.
From a random sample :
is the unbiased estimate of and is the unbiased estimate of . Both formulae are in the formula booklet.
If the variance of the sample (dividing by ) is already known, then
For large samples the factor is close to and the two versions are nearly equal. For small samples the difference is significant, and in every case the mark scheme expects .
The forms the data come in
Questions give the data in one of four ways. The method is always the same: get , and (or their coded versions), then substitute.
Raw data. Add up the values and their squares, or use the statistics mode of your calculator. Most calculators show two standard deviations: one labelled or (dividing by ) and one labelled or (dividing by ). Use the second, and square it.
Summary totals. You are given , and . Substitute directly.
Coded totals. You are given and for some constant . Subtracting a constant shifts every value by the same amount, which moves the mean but leaves the spread unchanged. So
A frequency table. Use and in place of and , with .
- Write down , and (or and , or and ).
- Mean: , adding back if the data are coded.
- Variance: , with coded totals in place of and if given. Never add to the variance.
- Give answers to 3 significant figures (or exactly), and keep more figures if the value is used later.
A random sample of packets of rice has masses, in grams,
Find unbiased estimates of the population mean and variance.
Solution
, , .
Check with deviations from : , whose squares add to . Dividing by gives . With large numbers like these, the deviation form avoids subtracting two huge, nearly equal totals.
(a) For a random sample of values of , and . Find unbiased estimates of the population mean and variance.
(b) The times, seconds, taken by a random sample of students to solve a puzzle are summarised by and . Find unbiased estimates of the population mean and variance of .
Solution
(a)
(b) The mean of the coded values is , so
The variance is unaffected by subtracting :
The numbers of goals scored in a random sample of football matches are shown.
| Goals, | ||||||
|---|---|---|---|---|---|---|
| Frequency, |
(a) Find unbiased estimates of the mean and variance of the number of goals per match.
(b) Comment on whether a Poisson distribution could be a suitable model.
Solution
(a)
(b) For a Poisson distribution the mean equals the variance. The estimates and are close, so a Poisson model is plausible (see modelling with the Poisson distribution).
A calculator gives the standard deviation of a sample of values, calculated by dividing by , as . Find the unbiased estimate of the population variance.
Solution
The variance of the sample is , so
A random sample of observations of a random variable has mean , and the unbiased estimate of the population variance from this sample is . A second, independent random sample of observations of gives and .
(a) Using both samples together, find unbiased estimates of the mean and variance of .
(b) Use your estimates to find the approximate probability that the mean of a further random sample of observations of is greater than .
Solution
(a) Recover the totals for the first sample.
From :
Combined: , , .
(b) The distribution of is unknown, but is large, so by the Central Limit Theorem
You cannot simply average the two variance estimates: the samples have different sizes and different means, and the spread of the combined data includes the gap between those means.
Where the estimates are used
When is unknown and the sample is large, stands in for in confidence intervals and in tests for a population mean. An error in therefore carries through the rest of a long question, which is why the method mark for this step is so valuable.
In the same spirit, the sample proportion is an unbiased estimate of a population proportion , because when . It is used in confidence intervals for proportions.
- Dividing by . The unbiased estimate of variance divides by . Using loses the accuracy mark, and everything built on it is wrong.
- Using the wrong calculator value. (or ) divides by ; (or ) divides by . Square the right one.
- Adding the coding constant to the variance. needs added back to get the mean; the variance needs nothing added.
- Mixing coded and uncoded totals. In both totals must be coded.
- Forgetting to square . It is , not .
- Giving when asked for the variance, or when a standard deviation is needed in the next part. Read which one is asked for.
- Explaining "unbiased" as "the sample is random" or "the estimate is accurate". It means the expected value of the estimator equals the parameter, so the method is right on average.
- The question words "unbiased estimates of the population mean and variance" signal exactly this calculation. Present it as two separate lines, formula then numbers.
- Show the substitution into the formula, such as . A slip after a correct substitution keeps the method mark.
- Give answers exactly where they terminate (, ) or to 3 significant figures. If the value is used later, carry more figures.
- When a later part needs the standard deviation, take from the unrounded value.
- "Explain what is meant by an unbiased estimate" is worth one mark: "the expected value of the estimator equals the population parameter, so on average it gives the true value".
- An estimator is unbiased if its expected value equals the parameter: right on average, though individual estimates vary.
- is the unbiased estimate of .
- is the unbiased estimate of .
- Dividing by underestimates on average, by the factor .
- (variance of the sample, dividing by ).
- For coded data, add back to the mean only; the variance is unchanged by coding.
- For frequency tables use , and .
Practice questions
-
A random sample gives the values . Find unbiased estimates of the population mean and variance.
-
For a random sample of values, and . Find unbiased estimates of the population mean and variance.
-
For a random sample of values, and . Find unbiased estimates of the population mean and variance.
-
The variance of a random sample of values, calculated by dividing by , is . Find the unbiased estimate of the population variance.
-
The numbers of children in a random sample of households are shown.
Children, Frequency Find unbiased estimates of the population mean and variance.
-
A researcher says that the sample mean is an unbiased estimator of the population mean. Explain what this means.
-
A random sample of observations has mean , and the unbiased estimate of the population variance is . A thirteenth observation, , is added to the sample. Find the new unbiased estimates of the population mean and variance.
-
For a random sample of values of , and , where is a constant. (a) Find the unbiased estimate of the population variance. (b) The unbiased estimate of the population mean is . Find .
Answers
-
, , . ; .
-
; .
-
; .
-
.
-
, , . ; .
-
The expected value of the sample mean equals the population mean, . Different samples give different sample means, some too high and some too low, but on average the sample mean equals the population mean.
-
Original: and . New: , , . ; .
-
(a) . Note that is not needed: coding does not affect the variance. (b) , so .