Type I and Type II Errors
A hypothesis test makes a decision from incomplete information, so it can be wrong. It can reject a null hypothesis that is actually true, or it can fail to reject one that is actually false. These two mistakes are called Type I and Type II errors. Paper 6 asks you to describe them in context, to say which one could have been made in a given test, and to calculate their probabilities for binomial, Poisson and normal tests. A Type I and Type II part ends a large share of the hypothesis-testing questions on the paper, and it is often where the hardest marks are.
Two ways to be wrong
A test has two possible decisions, and the truth has two possible states, so there are four combinations.
| is true | is false | |
|---|---|---|
| Reject | Type I error | Correct decision |
| Do not reject | Correct decision | Type II error |
- A Type I error is rejecting when is true.
- A Type II error is not rejecting (accepting) when is false.
A courtroom analogy helps: convicting an innocent person is a Type I error; acquitting a guilty one is a Type II error.
Because each error belongs to one decision, only one of them can have happened in any particular test. If was rejected, the only possible error is Type I. If was not rejected, the only possible error is Type II. This is a favourite one-mark question.
The probability of a Type I error
A Type I error happens when is true but the test statistic falls in the critical region. So
- For a test using a normal distribution (a test for a mean), the critical region is chosen to have probability exactly , so equals the significance level.
- For a test using binomial or Poisson probabilities, the critical region usually has probability less than . Then is the actual significance level: the exact probability of the critical region under . You must find the critical region first.
The probability of a Type II error
A Type II error happens when is false but the test statistic falls in the acceptance region. To calculate a probability you need to know the true value of the parameter, so a question will always give one: "find the probability of a Type II error if the true proportion is ".
The critical region is still the one found using . Only the distribution used to calculate the probability changes.
- Find the critical region using the distribution under (this may already have been done in an earlier part).
- : the probability of the critical region under . For a normal test this is just .
- For a Type II error, write down the distribution of the test statistic with the true parameter value given in the question.
- : the probability, under that true distribution, that the test statistic is not in the critical region (it lies in the acceptance region).
The graph shows both errors for a test of against , using the mean of bags with g, so that has standard deviation . The right-hand curve is the distribution of if is true; the left-hand curve is its distribution if in fact . The critical region is . The shaded area on the left, under the curve, is . The shaded area on the right, under the curve, is .
The picture shows the trade-off between the two errors. Moving the critical value to the left (a smaller significance level) shrinks the Type I area but enlarges the Type II area. The only way to reduce both at once is to make both curves narrower, by taking a larger sample.
It is claimed that of customers at a shop use a discount code. The manager suspects the proportion is lower and tests the claim at the significance level, using a random sample of customers.
(a) Find the critical region and the probability of a Type I error.
(b) Find the probability of a Type II error if the true proportion is .
Solution
(a) , . Under , .
The critical region is , and
(b) If , then . A Type II error occurs if is not in the critical region, that is .
The number of accidents per month at a junction has had the distribution . After new signs are installed, the council will test at the significance level whether the accident rate has fallen, using the number of accidents in one month.
(a) Find the critical region and the probability of a Type I error.
(b) Find the probability of a Type II error if the mean number of accidents per month is now .
Solution
(a) , . Under , .
The critical region is , and .
(b) If , then . A Type II error occurs if .
This is large: with only one month of data, even a big fall in the accident rate will usually go undetected. Using several months of data would reduce it.
Bags of sugar have masses grams, and should be . A random sample of bags is weighed and is tested against at the significance level.
(a) State the probability of a Type I error.
(b) Find the probability of a Type II error if the true mean is g.
Solution
(a) The test statistic is normal, so .
(b) Under , . Reject if
If , then . A Type II error occurs if .
In the sugar test above, a sample of bags has mean mass g.
(a) Carry out the test.
(b) State which type of error could have been made, and describe it in context.
Solution
(a) , so is in the critical region. Reject : there is evidence at the level that the mean mass of the bags is less than g.
(b) was rejected, so the only possible error is a Type I error: concluding that the mean mass of the bags is less than g when in fact it is g.
(A Type II error, here, would be concluding that there is no evidence the bags are underweight when in fact their mean mass is less than g. It could not have happened in this test, because was rejected.)
The marks in a national test are . A test of against at the significance level uses the mean of a random sample of marks.
(a) For , find the probability of a Type II error if in fact .
(b) Find the smallest value of for which the probability of a Type II error, when , is less than . You may neglect the lower tail of the critical region.
Solution
(a) Under , . The acceptance region is
If , , and
For a two-tailed test, the acceptance region has two boundaries, and the Type II probability is the area between them under the true distribution. Here the lower boundary is so far from that it contributes nothing.
(b) With sample size , the standard deviation of is , and the upper boundary of the acceptance region is .
If , we need
Simplify the left-hand side: , so
The smallest sample size is .
- Using to find a Type II probability. The critical region comes from , but the probability of a Type II error is calculated with the true parameter value.
- Saying for a binomial or Poisson test. It is the actual significance level, such as .
- Using the critical region instead of the acceptance region for a Type II error. Type II means is not rejected, so you need the probability of the acceptance region.
- Mixing up the definitions. Type I: reject a true . Type II: fail to reject a false . A useful memory aid: Type I is the error you control directly with the significance level.
- Saying both errors were possible. In a particular test only one is: Type I if was rejected, Type II if not.
- Forgetting the second boundary in a two-tailed Type II calculation. Check whether the far boundary contributes; usually it adds almost nothing, but say so.
- Describing errors without context. "Rejecting when it is true" earns less than "concluding the mean mass is less than g when it is g".
- "Explain what is meant by a Type I error in this context" wants the definition applied to the question: what would be concluded, and what is actually true.
- "State which type of error could have been made" needs the reason: "Type I, because was rejected".
- For binomial and Poisson tests, the Type I probability needs the critical region. If an earlier part used the probability method, you must now find the critical region.
- For a Type II probability, write the true distribution explicitly, such as "if , ", and state the event, such as "Type II error: ".
- Normal tests: the critical value of (such as ) is the key intermediate value. Keep at least 4 significant figures in it.
- The syllabus asks for error probabilities in tests based on a normal distribution or on direct evaluation of binomial or Poisson probabilities. Set the work out in the same way whichever of these the question uses: critical region from , then the probability under the stated true value.
- Type I error: rejecting when it is true. Type II error: not rejecting when it is false.
- : equal to for a normal test, the actual significance level for a discrete test.
- ; a specific true value is needed.
- Find the critical region using ; find the Type II probability using the true distribution.
- Only one error is possible in a given test: Type I if was rejected, Type II if not.
- Lowering reduces Type I but increases Type II; a larger sample reduces Type II for the same .
- Describe errors in the context of the question.
Practice questions
- A test of against is carried out at the significance level, using . Find the critical region and the probability of a Type I error.
- For the test in question 1, find the probability of a Type II error if in fact .
- A test of against is carried out at the significance level, using a single observation of . Find the probability of a Type II error if in fact .
- . A test of against at the significance level uses the mean of a random sample of observations. Find the probability of a Type II error if in fact .
- A drug company tests whether a new drug cures a higher proportion of patients than the old one. The test result is to reject . State which type of error could have been made, and describe it in context.
- (a) Explain why the probability of a Type II error cannot be found without a specific value of the parameter. (b) Explain why reducing the significance level of a test increases the probability of a Type II error.
- A test of against uses and has critical region or . Find the probability of a Type I error, and the probability of a Type II error if in fact .
- Breakdowns of a machine occur at random at an average rate of per week. After an overhaul, a manager will test at the significance level whether the rate has decreased, using the total number of breakdowns in weeks. (a) Find the critical region and the probability of a Type I error. (b) Find the probability of a Type II error if the true rate after the overhaul is per week. (c) The manager's assistant suggests using just week instead. Find the critical region and the probability of a Type II error in this case, and comment.
Answers
-
Under , . and . Critical region ; .
-
If , . Type II error: . .
-
Under , : and , so the critical region is . If , . Type II error: . .
-
Under , . Reject if . If , . .
-
was rejected, so a Type I error could have been made: concluding that the new drug cures a higher proportion of patients when in fact it cures the same proportion as the old one.
-
(a) A Type II error happens when is false, and its probability is the probability of the acceptance region under the true distribution. " is false" does not say what the true value is, and the probability depends on it (it is larger when the true value is close to the value), so a specific value is needed. (b) A smaller significance level makes the critical region smaller, so the acceptance region is larger. A larger acceptance region has a larger probability under the true distribution, so a Type II error becomes more likely.
-
. If , . Type II error: . .
-
(a) Over weeks, under , . and . Critical region ; . (b) True mean over weeks , so . Type II error: . . (c) For week, under , . and , so the critical region is . If the true rate is per week, , and . The one-week test is much more likely to miss a genuine decrease ( against ), so the three-week test is better. Longer observation periods give the test more power to detect a change.