Type I and Type II Errors

A2 · S2 · 16 min

A hypothesis test makes a decision from incomplete information, so it can be wrong. It can reject a null hypothesis that is actually true, or it can fail to reject one that is actually false. These two mistakes are called Type I and Type II errors. Paper 6 asks you to describe them in context, to say which one could have been made in a given test, and to calculate their probabilities for binomial, Poisson and normal tests. A Type I and Type II part ends a large share of the hypothesis-testing questions on the paper, and it is often where the hardest marks are.

Two ways to be wrong

A test has two possible decisions, and the truth has two possible states, so there are four combinations.

H0H_0 is trueH0H_0 is false
Reject H0H_0Type I errorCorrect decision
Do not reject H0H_0Correct decisionType II error
Definition
  • A Type I error is rejecting H0H_0 when H0H_0 is true.
  • A Type II error is not rejecting (accepting) H0H_0 when H0H_0 is false.

A courtroom analogy helps: convicting an innocent person is a Type I error; acquitting a guilty one is a Type II error.

Because each error belongs to one decision, only one of them can have happened in any particular test. If H0H_0 was rejected, the only possible error is Type I. If H0H_0 was not rejected, the only possible error is Type II. This is a favourite one-mark question.

The probability of a Type I error

A Type I error happens when H0H_0 is true but the test statistic falls in the critical region. So

P(Type I error)=P(test statistic in the critical region∣H0 true).P(\text{Type I error}) = P(\text{test statistic in the critical region} \mid H_0 \text{ true}).
  • For a test using a normal distribution (a test for a mean), the critical region is chosen to have probability exactly α\alpha, so P(Type I error)P(\text{Type I error}) equals the significance level.
  • For a test using binomial or Poisson probabilities, the critical region usually has probability less than α\alpha. Then P(Type I error)P(\text{Type I error}) is the actual significance level: the exact probability of the critical region under H0H_0. You must find the critical region first.

The probability of a Type II error

A Type II error happens when H0H_0 is false but the test statistic falls in the acceptance region. To calculate a probability you need to know the true value of the parameter, so a question will always give one: "find the probability of a Type II error if the true proportion is 0.10.1".

P(Type II error)=P(test statistic not in the critical region∣the given true value).P(\text{Type II error}) = P(\text{test statistic not in the critical region} \mid \text{the given true value}).

The critical region is still the one found using H0H_0. Only the distribution used to calculate the probability changes.

Probabilities of Type I and Type II errors
  1. Find the critical region using the distribution under H0H_0 (this may already have been done in an earlier part).
  2. P(Type I error)P(\text{Type I error}): the probability of the critical region under H0H_0. For a normal test this is just α\alpha.
  3. For a Type II error, write down the distribution of the test statistic with the true parameter value given in the question.
  4. P(Type II error)P(\text{Type II error}): the probability, under that true distribution, that the test statistic is not in the critical region (it lies in the acceptance region).

The graph shows both errors for a test of H0:μ=1000H_0: \mu = 1000 against H1:μ<1000H_1: \mu < 1000, using the mean of 1616 bags with σ=8\sigma = 8 g, so that Xˉ\bar{X} has standard deviation 22. The right-hand curve is the distribution of Xˉ\bar{X} if H0H_0 is true; the left-hand curve is its distribution if in fact μ=995\mu = 995. The critical region is xˉ<996.71\bar{x} < 996.71. The shaded area on the left, under the H0H_0 curve, is P(Type I error)=0.05P(\text{Type I error}) = 0.05. The shaded area on the right, under the μ=995\mu = 995 curve, is P(Type II error)P(\text{Type II error}).

y = exp(-(x - 1000)^2 / 8) / (2 sqrt(2 pi)) y = exp(-(x - 995)^2 / 8) / (2 sqrt(2 pi)) fill 988 996.71 y = exp(-(x - 1000)^2 / 8) / (2 sqrt(2 pi)) fill 996.71 1003 y = exp(-(x - 995)^2 / 8) / (2 sqrt(2 pi))

The picture shows the trade-off between the two errors. Moving the critical value to the left (a smaller significance level) shrinks the Type I area but enlarges the Type II area. The only way to reduce both at once is to make both curves narrower, by taking a larger sample.

Routine: a binomial test

It is claimed that 30%30\% of customers at a shop use a discount code. The manager suspects the proportion is lower and tests the claim at the 5%5\% significance level, using a random sample of 2020 customers.

(a) Find the critical region and the probability of a Type I error.

(b) Find the probability of a Type II error if the true proportion is 0.10.1.

Solution

(a) H0:p=0.3H_0: p = 0.3, H1:p<0.3\quad H_1: p < 0.3. Under H0H_0, X∼B(20,0.3)X \sim B(20, 0.3).

P(X≤2)=0.0355≤0.05,P(X≤3)=0.1071>0.05.P(X \le 2) = 0.0355 \le 0.05, \qquad P(X \le 3) = 0.1071 > 0.05.

The critical region is X≤2X \le 2, and

P(Type I error)=P(X≤2∣p=0.3)=0.0355.P(\text{Type I error}) = P(X \le 2 \mid p = 0.3) = 0.0355.

(b) If p=0.1p = 0.1, then X∼B(20,0.1)X \sim B(20, 0.1). A Type II error occurs if XX is not in the critical region, that is X≥3X \ge 3.

P(X≥3)=1−(0.920+20(0.1)(0.9)19+190(0.1)2(0.9)18)=1−(0.1216+0.2702+0.2852)=1−0.6769=0.323P(X \ge 3) = 1 - \left(0.9^{20} + 20(0.1)(0.9)^{19} + 190(0.1)^2(0.9)^{18}\right) = 1 - (0.1216 + 0.2702 + 0.2852) = 1 - 0.6769 = 0.323
A Poisson test

The number of accidents per month at a junction has had the distribution Po(4)\text{Po}(4). After new signs are installed, the council will test at the 5%5\% significance level whether the accident rate has fallen, using the number of accidents in one month.

(a) Find the critical region and the probability of a Type I error.

(b) Find the probability of a Type II error if the mean number of accidents per month is now 1.51.5.

Solution

(a) H0:λ=4H_0: \lambda = 4, H1:λ<4\quad H_1: \lambda < 4. Under H0H_0, X∼Po(4)X \sim \text{Po}(4).

P(X=0)=e−4=0.0183≤0.05,P(X≤1)=5e−4=0.0916>0.05.P(X = 0) = e^{-4} = 0.0183 \le 0.05, \qquad P(X \le 1) = 5e^{-4} = 0.0916 > 0.05.

The critical region is X=0X = 0, and P(Type I error)=0.0183P(\text{Type I error}) = 0.0183.

(b) If λ=1.5\lambda = 1.5, then X∼Po(1.5)X \sim \text{Po}(1.5). A Type II error occurs if X≥1X \ge 1.

P(X≥1)=1−e−1.5=0.777P(X \ge 1) = 1 - e^{-1.5} = 0.777

This is large: with only one month of data, even a big fall in the accident rate will usually go undetected. Using several months of data would reduce it.

A normal test

Bags of sugar have masses N(μ,82)N(\mu, 8^2) grams, and μ\mu should be 10001000. A random sample of 1616 bags is weighed and H0:μ=1000H_0: \mu = 1000 is tested against H1:μ<1000H_1: \mu < 1000 at the 5%5\% significance level.

(a) State the probability of a Type I error.

(b) Find the probability of a Type II error if the true mean is 995995 g.

Solution

(a) The test statistic is normal, so P(Type I error)=0.05P(\text{Type I error}) = 0.05.

(b) Under H0H_0, Xˉ∼N(1000,6416)=N(1000,4)\bar{X} \sim N\left(1000, \dfrac{64}{16}\right) = N(1000, 4). Reject H0H_0 if

xˉ<1000−1.645×2=996.71.\bar{x} < 1000 - 1.645 \times 2 = 996.71.

If μ=995\mu = 995, then Xˉ∼N(995,4)\bar{X} \sim N(995, 4). A Type II error occurs if Xˉ≥996.71\bar{X} \ge 996.71.

P(Xˉ≥996.71)=P(Z≥996.71−9952)=P(Z≥0.855)=1−0.8037=0.196P(\bar{X} \ge 996.71) = P\left(Z \ge \frac{996.71 - 995}{2}\right) = P(Z \ge 0.855) = 1 - 0.8037 = 0.196
Which error, in context

In the sugar test above, a sample of 1616 bags has mean mass 995.9995.9 g.

(a) Carry out the test.

(b) State which type of error could have been made, and describe it in context.

Solution

(a) 995.9<996.71995.9 < 996.71, so xˉ\bar{x} is in the critical region. Reject H0H_0: there is evidence at the 5%5\% level that the mean mass of the bags is less than 10001000 g.

(b) H0H_0 was rejected, so the only possible error is a Type I error: concluding that the mean mass of the bags is less than 10001000 g when in fact it is 10001000 g.

(A Type II error, here, would be concluding that there is no evidence the bags are underweight when in fact their mean mass is less than 10001000 g. It could not have happened in this test, because H0H_0 was rejected.)

Exam-hard: a two-tailed test and the sample size needed

The marks in a national test are N(μ,52)N(\mu, 5^2). A test of H0:μ=50H_0: \mu = 50 against H1:μ≠50H_1: \mu \ne 50 at the 5%5\% significance level uses the mean of a random sample of nn marks.

(a) For n=25n = 25, find the probability of a Type II error if in fact μ=53\mu = 53.

(b) Find the smallest value of nn for which the probability of a Type II error, when μ=53\mu = 53, is less than 0.010.01. You may neglect the lower tail of the critical region.

Solution

(a) Under H0H_0, Xˉ∼N(50,1)\bar{X} \sim N(50, 1). The acceptance region is

50−1.96<xˉ<50+1.96,that is48.04<xˉ<51.96.50 - 1.96 < \bar{x} < 50 + 1.96, \quad\text{that is}\quad 48.04 < \bar{x} < 51.96.

If μ=53\mu = 53, Xˉ∼N(53,1)\bar{X} \sim N(53, 1), and

P(48.04<Xˉ<51.96)=P(−4.96<Z<−1.04)=(1−0.8508)−0.0000=0.149.P(48.04 < \bar{X} < 51.96) = P(-4.96 < Z < -1.04) = (1 - 0.8508) - 0.0000 = 0.149.

For a two-tailed test, the acceptance region has two boundaries, and the Type II probability is the area between them under the true distribution. Here the lower boundary is so far from 5353 that it contributes nothing.

(b) With sample size nn, the standard deviation of Xˉ\bar{X} is 5n\dfrac{5}{\sqrt{n}}, and the upper boundary of the acceptance region is 50+1.96×5n50 + 1.96 \times \dfrac{5}{\sqrt{n}}.

If μ=53\mu = 53, we need

P(Xˉ<50+9.8n)<0.01⟺50+9.8/n−535/n<−2.326.P\left(\bar{X} < 50 + \frac{9.8}{\sqrt{n}}\right) < 0.01 \quad\Longleftrightarrow\quad \frac{50 + 9.8/\sqrt{n} - 53}{5/\sqrt{n}} < -2.326.

Simplify the left-hand side: 1.96−3n5<−2.3261.96 - \dfrac{3\sqrt{n}}{5} < -2.326, so

3n5>4.286  ⇒  n>7.143  ⇒  n>51.03.\frac{3\sqrt{n}}{5} > 4.286 \;\Rightarrow\; \sqrt{n} > 7.143 \;\Rightarrow\; n > 51.03.

The smallest sample size is n=52n = 52.

Common mistakes
  • Using H0H_0 to find a Type II probability. The critical region comes from H0H_0, but the probability of a Type II error is calculated with the true parameter value.
  • Saying P(Type I error)=5%P(\text{Type I error}) = 5\% for a binomial or Poisson test. It is the actual significance level, such as 0.03550.0355.
  • Using the critical region instead of the acceptance region for a Type II error. Type II means H0H_0 is not rejected, so you need the probability of the acceptance region.
  • Mixing up the definitions. Type I: reject a true H0H_0. Type II: fail to reject a false H0H_0. A useful memory aid: Type I is the error you control directly with the significance level.
  • Saying both errors were possible. In a particular test only one is: Type I if H0H_0 was rejected, Type II if not.
  • Forgetting the second boundary in a two-tailed Type II calculation. Check whether the far boundary contributes; usually it adds almost nothing, but say so.
  • Describing errors without context. "Rejecting H0H_0 when it is true" earns less than "concluding the mean mass is less than 10001000 g when it is 10001000 g".
Exam tip
  • "Explain what is meant by a Type I error in this context" wants the definition applied to the question: what would be concluded, and what is actually true.
  • "State which type of error could have been made" needs the reason: "Type I, because H0H_0 was rejected".
  • For binomial and Poisson tests, the Type I probability needs the critical region. If an earlier part used the probability method, you must now find the critical region.
  • For a Type II probability, write the true distribution explicitly, such as "if p=0.1p = 0.1, X∼B(20,0.1)X \sim B(20, 0.1)", and state the event, such as "Type II error: X≥3X \ge 3".
  • Normal tests: the critical value of xˉ\bar{x} (such as 996.71996.71) is the key intermediate value. Keep at least 4 significant figures in it.
  • The syllabus asks for error probabilities in tests based on a normal distribution or on direct evaluation of binomial or Poisson probabilities. Set the work out in the same way whichever of these the question uses: critical region from H0H_0, then the probability under the stated true value.
Summary
  • Type I error: rejecting H0H_0 when it is true. Type II error: not rejecting H0H_0 when it is false.
  • P(Type I)=P(critical region∣H0)P(\text{Type I}) = P(\text{critical region} \mid H_0): equal to α\alpha for a normal test, the actual significance level for a discrete test.
  • P(Type II)=P(acceptance region∣true value)P(\text{Type II}) = P(\text{acceptance region} \mid \text{true value}); a specific true value is needed.
  • Find the critical region using H0H_0; find the Type II probability using the true distribution.
  • Only one error is possible in a given test: Type I if H0H_0 was rejected, Type II if not.
  • Lowering α\alpha reduces Type I but increases Type II; a larger sample reduces Type II for the same α\alpha.
  • Describe errors in the context of the question.

Practice questions

Question
  1. A test of H0:p=0.4H_0: p = 0.4 against H1:p>0.4H_1: p > 0.4 is carried out at the 5%5\% significance level, using X∼B(15,p)X \sim B(15, p). Find the critical region and the probability of a Type I error.
  2. For the test in question 1, find the probability of a Type II error if in fact p=0.7p = 0.7.
  3. A test of H0:λ=6H_0: \lambda = 6 against H1:λ>6H_1: \lambda > 6 is carried out at the 5%5\% significance level, using a single observation of X∼Po(λ)X \sim \text{Po}(\lambda). Find the probability of a Type II error if in fact λ=10\lambda = 10.
  4. X∼N(μ,52)X \sim N(\mu, 5^2). A test of H0:μ=50H_0: \mu = 50 against H1:μ>50H_1: \mu > 50 at the 5%5\% significance level uses the mean of a random sample of 2525 observations. Find the probability of a Type II error if in fact μ=53\mu = 53.
  5. A drug company tests whether a new drug cures a higher proportion of patients than the old one. The test result is to reject H0H_0. State which type of error could have been made, and describe it in context.
  6. (a) Explain why the probability of a Type II error cannot be found without a specific value of the parameter. (b) Explain why reducing the significance level of a test increases the probability of a Type II error.
  7. A test of H0:p=0.3H_0: p = 0.3 against H1:p≠0.3H_1: p \ne 0.3 uses X∼B(20,p)X \sim B(20, p) and has critical region X≤1X \le 1 or X≥11X \ge 11. Find the probability of a Type I error, and the probability of a Type II error if in fact p=0.5p = 0.5.
  8. Breakdowns of a machine occur at random at an average rate of 33 per week. After an overhaul, a manager will test at the 5%5\% significance level whether the rate has decreased, using the total number of breakdowns in 33 weeks. (a) Find the critical region and the probability of a Type I error. (b) Find the probability of a Type II error if the true rate after the overhaul is 1.51.5 per week. (c) The manager's assistant suggests using just 11 week instead. Find the critical region and the probability of a Type II error in this case, and comment.
Answers
  1. Under H0H_0, X∼B(15,0.4)X \sim B(15, 0.4). P(X≥10)=0.0338≤0.05P(X \ge 10) = 0.0338 \le 0.05 and P(X≥9)=0.0950>0.05P(X \ge 9) = 0.0950 > 0.05. Critical region X≥10X \ge 10; P(Type I error)=0.0338P(\text{Type I error}) = 0.0338.

  2. If p=0.7p = 0.7, X∼B(15,0.7)X \sim B(15, 0.7). Type II error: X≤9X \le 9. P(X≤9)=0.278P(X \le 9) = 0.278.

  3. Under H0H_0, X∼Po(6)X \sim \text{Po}(6): P(X≥11)=0.0426≤0.05P(X \ge 11) = 0.0426 \le 0.05 and P(X≥10)=0.0839>0.05P(X \ge 10) = 0.0839 > 0.05, so the critical region is X≥11X \ge 11. If λ=10\lambda = 10, X∼Po(10)X \sim \text{Po}(10). Type II error: X≤10X \le 10. P(X≤10)=0.583P(X \le 10) = 0.583.

  4. Under H0H_0, Xˉ∼N(50,1)\bar{X} \sim N(50, 1). Reject if xˉ>50+1.645=51.645\bar{x} > 50 + 1.645 = 51.645. If μ=53\mu = 53, Xˉ∼N(53,1)\bar{X} \sim N(53, 1). P(Xˉ<51.645)=P(Z<−1.355)=1−0.9123=0.0877P(\bar{X} < 51.645) = P(Z < -1.355) = 1 - 0.9123 = 0.0877.

  5. H0H_0 was rejected, so a Type I error could have been made: concluding that the new drug cures a higher proportion of patients when in fact it cures the same proportion as the old one.

  6. (a) A Type II error happens when H0H_0 is false, and its probability is the probability of the acceptance region under the true distribution. "H0H_0 is false" does not say what the true value is, and the probability depends on it (it is larger when the true value is close to the H0H_0 value), so a specific value is needed. (b) A smaller significance level makes the critical region smaller, so the acceptance region is larger. A larger acceptance region has a larger probability under the true distribution, so a Type II error becomes more likely.

  7. P(Type I)=P(X≤1∣p=0.3)+P(X≥11∣p=0.3)=0.0076+0.0171=0.0248P(\text{Type I}) = P(X \le 1 \mid p = 0.3) + P(X \ge 11 \mid p = 0.3) = 0.0076 + 0.0171 = 0.0248. If p=0.5p = 0.5, X∼B(20,0.5)X \sim B(20, 0.5). Type II error: 2≤X≤102 \le X \le 10. P(2≤X≤10)=P(X≤10)−P(X≤1)=0.5881−0.0000=0.588P(2 \le X \le 10) = P(X \le 10) - P(X \le 1) = 0.5881 - 0.0000 = 0.588.

  8. (a) Over 33 weeks, under H0H_0, X∼Po(9)X \sim \text{Po}(9). P(X≤3)=0.0212≤0.05P(X \le 3) = 0.0212 \le 0.05 and P(X≤4)=0.0550>0.05P(X \le 4) = 0.0550 > 0.05. Critical region X≤3X \le 3; P(Type I error)=0.0212P(\text{Type I error}) = 0.0212. (b) True mean over 33 weeks =4.5= 4.5, so X∼Po(4.5)X \sim \text{Po}(4.5). Type II error: X≥4X \ge 4. P(X≥4)=1−e−4.5(1+4.5+4.522+4.536)=1−0.3423=0.658P(X \ge 4) = 1 - e^{-4.5}\left(1 + 4.5 + \tfrac{4.5^2}{2} + \tfrac{4.5^3}{6}\right) = 1 - 0.3423 = 0.658. (c) For 11 week, under H0H_0, X∼Po(3)X \sim \text{Po}(3). P(X=0)=e−3=0.0498≤0.05P(X = 0) = e^{-3} = 0.0498 \le 0.05 and P(X≤1)=4e−3=0.199>0.05P(X \le 1) = 4e^{-3} = 0.199 > 0.05, so the critical region is X=0X = 0. If the true rate is 1.51.5 per week, X∼Po(1.5)X \sim \text{Po}(1.5), and P(Type II)=P(X≥1)=1−e−1.5=0.777P(\text{Type II}) = P(X \ge 1) = 1 - e^{-1.5} = 0.777. The one-week test is much more likely to miss a genuine decrease (0.7770.777 against 0.6580.658), so the three-week test is better. Longer observation periods give the test more power to detect a change.

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action