Hypothesis Tests for a Population Mean
Is the mean mass of bags of sugar really g? Has a new training programme reduced the mean time to complete a task? Claims about a population mean are tested using the mean of a random sample. Under the null hypothesis the sample mean has a normal distribution, either exactly (normal population, known variance) or approximately (large sample, by the Central Limit Theorem), so the test is a -test using the normal tables. This is one of the most frequently examined tests on Paper 6, often combined with unbiased estimates from summarised data and followed by Type I and Type II errors.
The test statistic
To test , take a random sample of size and calculate . If is true, then from the distribution of the sample mean
and the standardised test statistic is
A large positive is evidence that ; a large negative is evidence that .
The syllabus covers the same two situations as for confidence intervals.
| Situation | Distribution of under | What to say |
|---|---|---|
| Population normal, known | exactly, any | " is normal, so is normal" |
| Large sample, any population | approximately; use if is unknown | " is large, so by the Central Limit Theorem is approximately normal" |
In the second case, is the unbiased estimate of .
| Significance level | |||||
|---|---|---|---|---|---|
| One-tailed | |||||
| Two-tailed |
For a lower-tail test, the critical value is negative: reject if at , for example. For a two-tailed test, reject if exceeds the critical value.
These values are printed under the normal table in the formula booklet. There is no continuity correction in a test for a mean: is already continuous (or treated as such through the Central Limit Theorem).
Two equivalent decisions
As with every test, you can decide in two ways.
Compare with the critical value. This is the most common approach for a mean.
Compare a probability with the significance level. For a lower-tail test, find and compare it with ; for an upper-tail test, .
The critical region for
Sometimes you are asked for the critical region in terms of the sample mean itself, for example before a sample is taken, or as the first step in finding the probability of a Type II error. Rearrange the critical -value:
The graph shows , the distribution of the mean mass of bags when is true and g. For a lower-tail test at , the critical region is , shaded.
A two-tailed test at the level rejects exactly when lies outside the confidence interval for . The two ideas are two views of the same calculation.
- Define in context and state and .
- If is unknown, find and from the data.
- State the distribution of under , with the justification (normal population, or large sample and the Central Limit Theorem).
- Calculate .
- Compare with the critical value (or compare the tail probability with ), showing the inequality.
- Decide about and conclude in context, without certainty.
A machine cuts rods whose lengths are normally distributed with mean mm and standard deviation mm. After the machine is serviced, a random sample of rods has mean length mm. Assuming the standard deviation is unchanged, test at the significance level whether the mean length has changed.
Solution
Let mm be the population mean length after the service.
, .
The population is normal, so under , .
The two-tailed critical value at is . Since , reject .
There is evidence at the level that the mean length of the rods has changed.
A manufacturer claims that the mean lifetime of its batteries is hours. A consumer group suspects that the mean is lower. The lifetimes, hours, of a random sample of batteries are summarised by
Test the consumer group's suspicion at the significance level.
Solution
Let hours be the population mean lifetime. , .
The distribution of lifetimes is not known, but is large, so by the Central Limit Theorem, under ,
The one-tailed critical value at is . Since , reject .
There is evidence at the level that the mean lifetime of the batteries is less than hours.
Bags of sugar have masses that are normally distributed with standard deviation g. The mean is supposed to be g. An inspector will weigh a random sample of bags and test at the significance level whether the mean is less than g.
(a) Find the critical region for the sample mean.
(b) The sample mean is g. State the conclusion.
Solution
(a) , . Under , .
Reject if
The critical region is .
(b) , so is not in the critical region. Do not reject : there is insufficient evidence at the level that the mean mass of the bags is less than g.
The marks in a national test are normally distributed with standard deviation . A teacher will take a random sample of students to test, at the significance level, whether the population mean differs from . Find the set of values of for which would be rejected.
Solution
, . Under , .
For a two-tailed test at , the critical values are . Reject if
A company states that the mean time taken to assemble a desk is minutes. A new set of instructions is introduced, and the times, minutes, for a random sample of customers are summarised by
(a) Find unbiased estimates of the population mean and variance of the assembly time.
(b) Test at the significance level whether the new instructions have reduced the mean assembly time.
(c) Find the smallest significance level, to the nearest , at which the conclusion in (b) would be reached, and state whether the conclusion would be the same at the level.
(d) Explain whether it was necessary to use the Central Limit Theorem.
Solution
(a)
(b) Let minutes be the population mean assembly time. , .
Under , approximately, with standard deviation .
, so reject . There is evidence at the level that the new instructions have reduced the mean assembly time.
(c) . is rejected at any level above , so the smallest level is .
At the critical value is , and , so would not be rejected: the conclusion would be different.
(d) Yes. The distribution of assembly times is not known (times are often skewed), so the Central Limit Theorem is needed to say that the sample mean is approximately normal. This is valid because is large. The theorem also justifies using in place of the unknown , since the sample is large.
- Using instead of in . The test is about the sample mean; standardise with its standard deviation.
- Hypotheses about . "" is wrong. The hypotheses are about the population mean .
- Wrong critical value. for one-tailed , for two-tailed . Check the number of tails before looking up the value.
- Sign errors in a lower-tail test. For the critical value is negative, and you reject if is more negative than it: .
- Dividing by in . Use the unbiased estimate.
- Adding a continuity correction. There is none in a test for a mean.
- Quoting the Central Limit Theorem for a normal population. If is normal, is exactly normal; the theorem is not needed.
- Typical mark allocation: hypotheses (1), correct distribution or standard deviation (1), correct (1), comparison with the correct critical value (1), conclusion in context (1). With unbiased estimates first, add one or two marks.
- Write the comparison explicitly: "". If you use probabilities instead, compare like with like: "".
- Keep to 3 or more significant figures. When is close to the critical value, a rounding slip can flip the conclusion.
- "State an assumption" questions: the standard deviation is unchanged by the new process; the sample is random; or the population is normal. Choose the one relevant to the question.
- "Is the Central Limit Theorem needed?" Yes if the population is not known to be normal (it is then needed because the sample is large enough); no if the population is normal.
- Conclusions use "evidence" language and refer to the context: "the mean lifetime of the batteries", not just "".
- Under , : exactly for a normal population, approximately for a large sample (Central Limit Theorem).
- Test statistic ; use for when unknown and is large.
- Critical values: (one-tailed ), (two-tailed ), (one-tailed ), (two-tailed ).
- Critical region for : in the appropriate direction.
- No continuity correction.
- A two-tailed test at rejects exactly when lies outside the confidence interval.
- Conclude in context, without certainty.
Practice questions
- The masses of bags of rice are normally distributed with standard deviation g. A random sample of bags has mean mass g. Test at the significance level whether the population mean is less than g.
- Scores on an aptitude test are normally distributed with mean and standard deviation . A random sample of candidates from a new school has mean score . Test at the significance level whether the mean score of candidates from this school differs from .
- For a random sample of values of a variable , and . Test at the significance level whether the population mean is greater than .
- . A test of against at the significance level uses the mean of a random sample of observations. Find the critical region for .
- A population has standard deviation . A test of against is carried out at the significance level, and the sample mean is . Find the smallest sample size for which would be rejected.
- State, with a reason, whether the Central Limit Theorem is needed in (a) question 2, (b) question 3.
- In the desk-assembly example above, explain what is meant by "the significance level" in the context of that test.
- A test of against is carried out at the significance level, using the mean of a random sample of observations from a normal population with standard deviation . (a) Find the set of values of for which is rejected. (b) Two independent samples of are taken and the test is carried out on each. Given that the population mean really is , find the probability that is rejected in exactly one of the two tests.
Answers
-
, . Under , , standard deviation . . Critical value ; . Do not reject : there is insufficient evidence at the level that the mean mass is less than g.
-
, . Under , , standard deviation . . Reject : there is evidence at the level that the mean score for this school differs from .
-
, , . , . By the Central Limit Theorem, approximately. . Do not reject : insufficient evidence at the level that the population mean is greater than .
-
Under , . Reject if or .
-
Need , so , . The smallest sample size is . (For this to be valid, either the population is normal or counts as large enough for the Central Limit Theorem.)
-
(a) No: the scores are normally distributed, so is exactly normal for any sample size. (b) Yes: the distribution of is not known, so the theorem is needed to treat as approximately normal; this is valid because is large.
-
If the new instructions really made no difference (mean minutes), there would be a probability of of concluding that they had reduced the mean assembly time.
-
(a) Under , , standard deviation . Reject if (to 4 s.f.; more precisely ). (b) Each test rejects a true with probability , independently. .