The Nature of Hypothesis Testing

A2 · S2 · 18 min

A hypothesis test is a formal way of deciding whether data give real evidence against a claim, or whether what you observed could easily have happened by chance. A manufacturer claims that only 5%5\% of its components are faulty; a sample has 99 faulty out of 8080. Is the claim false, or was this an unlucky sample? Every S2 test answers a question like this with the same logic and the same written structure, and every Paper 6 has at least one full test, usually worth 55 or 66 marks. This note sets out the language and the logic; the notes that follow apply them to binomial, Poisson and normal tests.

The logic of a test

Think of a courtroom. The defendant is presumed innocent, and is convicted only if the evidence would be very unlikely if they were innocent. A hypothesis test works the same way.

  1. Start by assuming the claim is true. This is the null hypothesis.
  2. Ask how likely a result at least as extreme as the one observed would be, if the claim were true.
  3. If that probability is very small (below a threshold fixed in advance), the data are hard to explain under the claim, so reject it. Otherwise, the data are consistent with the claim and you do not reject it.

Here is the whole idea in one small example. A coin is claimed to be fair, but you suspect it favours heads. You toss it 1010 times and get 99 heads. If the coin really is fair, the number of heads XX satisfies X∼B(10,0.5)X \sim B(10, 0.5), and

P(X≥9)=(109)+(1010)210=111024=0.0107.P(X \ge 9) = \frac{\binom{10}{9} + \binom{10}{10}}{2^{10}} = \frac{11}{1024} = 0.0107.

Nine or more heads would happen only about 1%1\% of the time with a fair coin. That is strong evidence that the coin is biased towards heads. It is not proof: fair coins do sometimes give 99 heads in 1010 tosses. A test controls how often this kind of wrong conclusion happens, but cannot rule it out.

The vocabulary

The syllabus lists the terms you must understand and use. Learn them precisely; definitions are sometimes asked for directly.

Definition
  • The null hypothesis, H0H_0, is the claim being tested, assumed true until the evidence says otherwise. It always states a single value of a population parameter, such as H0:p=0.5H_0: p = 0.5, H0:λ=6H_0: \lambda = 6 or H0:μ=500H_0: \mu = 500.
  • The alternative hypothesis, H1H_1, says how the parameter differs from that value if H0H_0 is false: H1:p>0.5H_1: p > 0.5, H1:p<0.5H_1: p < 0.5 or H1:p≠0.5H_1: p \ne 0.5.
  • A one-tailed test has an alternative hypothesis in one direction (>> or <<). A two-tailed test has H1H_1 with ≠\ne, so extreme results in either direction count as evidence.
  • The test statistic is the quantity calculated from the sample and used to decide: the number of successes, the number of events, or the sample mean (often converted to a zz-value).
  • The significance level is the threshold probability for rejecting H0H_0, such as 5%5\% or 1%1\%. It is the probability of rejecting H0H_0 when H0H_0 is true (for a discrete test statistic, the largest such probability allowed).
  • The critical region (or rejection region) is the set of values of the test statistic for which H0H_0 is rejected. The acceptance region is the set of values for which H0H_0 is not rejected. The critical value is the boundary value of the critical region.

Writing the hypotheses

Hypotheses are always about the population parameter, never about the sample. "H0:p=0.1H_0: p = 0.1" is about the proportion of the whole population; the sample result, such as 22 out of 5050, is the evidence.

The question's wording tells you which alternative to use.

Wording in the questionAlternative hypothesisTails
"has increased", "is more than", "is biased towards", "underestimates"H1:p>p0H_1: p > p_0One (upper)
"has decreased", "is less than", "has improved" (when lower is better)H1:p<p0H_1: p < p_0One (lower)
"has changed", "is different", "is biased" (no direction)H1:p≠p0H_1: p \ne p_0Two

The direction must be decided by the question, before looking at the data. Choosing a one-tailed test because the sample happens to be high is not allowed.

Two ways to reach a decision

Once H0H_0, H1H_1 and the distribution of the test statistic under H0H_0 are set out, there are two equivalent ways to decide.

Compare a probability with the significance level. Calculate the probability, assuming H0H_0, of a result at least as extreme as the one observed, in the direction of H1H_1. If it is less than the significance level, reject H0H_0. (This probability is often called the pp-value, though Paper 6 does not require the term.) For a two-tailed test, compare the probability in the observed tail with half the significance level.

Find the critical region. Work out which values of the test statistic would lead to rejection, then see whether the observed value is among them. This is essential when a question asks for the critical region, or for the probability of a Type I or Type II error.

For a test based on a normal distribution, the critical region is in one or both tails of the curve. The first graph shows a two-tailed test at the 5%5\% level, with 2.5%2.5\% in each tail beyond z=±1.96z = \pm 1.96; the second shows a one-tailed (upper) test at the 5%5\% level, with the whole 5%5\% beyond z=1.645z = 1.645.

y = exp(-x^2 / 2) / sqrt(2 pi) fill -4 -1.96 y = exp(-x^2 / 2) / sqrt(2 pi) fill 1.96 4 y = exp(-x^2 / 2) / sqrt(2 pi)
y = exp(-x^2 / 2) / sqrt(2 pi) fill 1.645 4 y = exp(-x^2 / 2) / sqrt(2 pi)

Discrete test statistics

For a binomial or Poisson test statistic, probabilities jump in steps, so a critical region usually cannot have probability exactly 5%5\%. The critical region is chosen to have probability as close as possible to, but not more than, the significance level. Its actual probability is called the actual significance level of the test. You will see this in detail in tests for a binomial proportion.

Writing the conclusion

The conclusion has two parts: the decision about H0H_0, and what it means in the context of the question. Both are needed for the final mark.

  • If the result is in the critical region: "Reject H0H_0. There is evidence at the 5%5\% level that the coin is biased towards heads."
  • If not: "Do not reject H0H_0. There is insufficient evidence at the 5%5\% level that the coin is biased towards heads." Cambridge mark schemes also accept "accept H0H_0", provided the contextual statement is not definite.

Never claim certainty. A test does not prove that H0H_0 is false, and failing to reject H0H_0 does not show that it is true; it only means the data are consistent with it. Avoid words like "proves", "shows that" and "the coin is fair".

Which test, which distribution

SituationTest statisticDistribution under H0H_0Note
Proportion, single observation of a count of successesNumber of successes XXB(n,p0)B(n, p_0)Binomial proportion
Rate of random eventsNumber of events XXPo(λ0)\text{Po}(\lambda_0)Poisson mean
Large nn or large λ\lambdaXX, with continuity correctionNormal approximationNormal approximations
Mean of a populationSample mean Xˉ\bar{X}N(μ0,σ2n)N\left(\mu_0, \dfrac{\sigma^2}{n}\right)Population mean
The structure of every hypothesis test
  1. Define the parameter in words and state H0H_0 and H1H_1 in symbols: "pp = probability of a head; H0:p=0.5H_0: p = 0.5, H1:p>0.5H_1: p > 0.5".
  2. State the test statistic and its distribution assuming H0H_0 is true: "X∼B(10,0.5)X \sim B(10, 0.5)".
  3. Calculate the probability of a result at least as extreme as the one observed (or find the critical region).
  4. Compare explicitly with the significance level (or say whether the observed value is in the critical region): "0.0107<0.050.0107 < 0.05".
  5. State the decision about H0H_0.
  6. Write the conclusion in context, without certainty.
Routine: the coin test in full

A coin is tossed 1010 times and shows 99 heads. Test at the 5%5\% significance level whether the coin is biased towards heads.

Solution

Let pp be the probability that the coin shows a head.

H0:p=0.5H_0: p = 0.5, H1:p>0.5\quad H_1: p > 0.5.

Let XX be the number of heads in 1010 tosses. Under H0H_0, X∼B(10,0.5)X \sim B(10, 0.5).

P(X≥9)=(109)(0.5)10+(0.5)10=111024=0.0107P(X \ge 9) = \binom{10}{9}(0.5)^{10} + (0.5)^{10} = \frac{11}{1024} = 0.0107

0.0107<0.050.0107 < 0.05, so reject H0H_0.

There is evidence at the 5%5\% significance level that the coin is biased towards heads.

Choosing the hypotheses

For each situation, define a parameter, and state suitable null and alternative hypotheses. Say whether the test is one-tailed or two-tailed.

(a) A website claims that 30%30\% of visitors buy something. A manager believes the true figure is lower.

(b) Emails arrive at an average rate of 66 per hour. After a new spam filter is installed, an analyst wants to know whether the rate has changed.

(c) The mean mass of bags of flour is supposed to be 10001000 g. A trading standards officer suspects that the bags are underweight on average.

Solution

(a) pp = proportion of all visitors who buy something. H0:p=0.3H_0: p = 0.3, H1:p<0.3H_1: p < 0.3. One-tailed.

(b) λ\lambda = mean number of emails per hour. H0:λ=6H_0: \lambda = 6, H1:λ≠6H_1: \lambda \ne 6. Two-tailed, because "changed" has no direction.

(c) μ\mu = population mean mass of bags, in grams. H0:μ=1000H_0: \mu = 1000, H1:μ<1000H_1: \mu < 1000. One-tailed.

In each case H0H_0 gives a single value, so the distribution of the test statistic can be written down exactly.

A two-tailed test: halve the level

A die is thrown 3030 times and a six appears only once. Test at the 5%5\% significance level whether the die is biased.

Solution

Let pp be the probability of a six. H0:p=16H_0: p = \tfrac{1}{6}, H1:p≠16\quad H_1: p \ne \tfrac{1}{6} (two-tailed).

Let XX be the number of sixes in 3030 throws. Under H0H_0, X∼B(30,16)X \sim B\left(30, \tfrac{1}{6}\right). The observed value 11 is below the expected value 55, so look at the lower tail.

P(X≤1)=(56)30+30(16)(56)29=0.00421+0.02528=0.0295P(X \le 1) = \left(\tfrac{5}{6}\right)^{30} + 30\left(\tfrac{1}{6}\right)\left(\tfrac{5}{6}\right)^{29} = 0.00421 + 0.02528 = 0.0295

The test is two-tailed, so compare with 2.5%2.5\%: 0.0295>0.0250.0295 > 0.025. Do not reject H0H_0.

There is insufficient evidence at the 5%5\% level that the die is biased.

Notice that a one-tailed test of H1:p<16H_1: p < \tfrac{1}{6} would have rejected H0H_0, since 0.0295<0.050.0295 < 0.05. The choice of tails matters, which is why it must be made from the question, not from the data.

Two equivalent methods for a normal test

The volumes of drink in bottles are normally distributed with standard deviation 22 ml. The mean is supposed to be 500500 ml. A random sample of 2525 bottles has mean volume 499.2499.2 ml. Test at the 5%5\% significance level whether the mean volume has changed.

Solution

Let μ\mu be the population mean volume in ml. H0:μ=500H_0: \mu = 500, H1:μ≠500\quad H_1: \mu \ne 500 (two-tailed).

Under H0H_0, Xˉ∼N(500,2225)\bar{X} \sim N\left(500, \dfrac{2^2}{25}\right), standard deviation 0.40.4.

Method 1, comparing zz with the critical value.

z=499.2−5000.4=−2.0z = \frac{499.2 - 500}{0.4} = -2.0

The critical values for a two-tailed test at 5%5\% are ±1.96\pm 1.96. Since −2.0<−1.96-2.0 < -1.96, the result is in the critical region.

Method 2, comparing a probability with the level.

P(Xˉ≤499.2)=P(Z≤−2.0)=1−0.9772=0.0228<0.025.P(\bar{X} \le 499.2) = P(Z \le -2.0) = 1 - 0.9772 = 0.0228 < 0.025.

Either way, reject H0H_0. There is evidence at the 5%5\% level that the mean volume of drink in the bottles has changed.

Exam-hard: the same data at different levels

It is known that 30%30\% of the members of a sports club are under 1818. After a recruitment drive, a random sample of 2525 members contains 1212 who are under 1818.

(a) Test at the 5%5\% significance level whether the proportion of members under 1818 has increased.

(b) Find the smallest significance level, correct to 1 decimal place, at which the test in (a) would lead to rejecting H0H_0.

(c) A colleague instead tests whether the proportion has changed, at the 5%5\% level. State, with a reason, the conclusion of this test.

Solution

(a) Let pp be the proportion of members who are under 1818.

H0:p=0.3H_0: p = 0.3, H1:p>0.3\quad H_1: p > 0.3.

Under H0H_0, X∼B(25,0.3)X \sim B(25, 0.3), where XX is the number under 1818 in the sample.

P(X≥12)=1−P(X≤11)=1−0.9558=0.0442P(X \ge 12) = 1 - P(X \le 11) = 1 - 0.9558 = 0.0442

(The sum P(X≤11)P(X \le 11) has twelve terms; in practice you would use a calculator's binomial cumulative function, and the method mark is for the expression 1−P(X≤11)1 - P(X \le 11) with B(25,0.3)B(25, 0.3).)

0.0442<0.050.0442 < 0.05, so reject H0H_0. There is evidence at the 5%5\% level that the proportion of members under 1818 has increased.

(b) H0H_0 is rejected whenever the significance level exceeds P(X≥12)=0.0442P(X \ge 12) = 0.0442, so the smallest level is 4.4%4.4\%.

(c) For a two-tailed test, the upper tail probability 0.04420.0442 is compared with 0.0250.025. Since 0.0442>0.0250.0442 > 0.025, H0H_0 is not rejected: there is insufficient evidence at the 5%5\% level that the proportion has changed.

Common mistakes
  • Hypotheses about the sample. "H0:xˉ=499.2H_0: \bar{x} = 499.2" or "H0:p=0.36H_0: p = 0.36" (the sample proportion) is wrong. Hypotheses are about the population parameter, with the value from the claim.
  • Inequalities in H0H_0. H0H_0 is always an equality. "H0:p≥0.5H_0: p \ge 0.5" is not accepted.
  • Comparing the wrong probability. Use the probability of the observed value or more extreme, never P(X=x)P(X = x) on its own. For "increased", that is P(X≥x)P(X \ge x); for "decreased", P(X≤x)P(X \le x).
  • Forgetting to halve the level in a two-tailed test. At the 5%5\% level, each tail gets 2.5%2.5\%.
  • Definite conclusions. "The coin is biased" or "this proves the mean has changed" loses the mark. Use "there is evidence that" or "there is insufficient evidence that".
  • No context. "Reject H0H_0" alone does not earn the final mark; say what it means in the situation.
  • Choosing the tail after seeing the data. The direction of H1H_1 comes from the question's wording.
Exam tip
  • Define the parameter in words ("pp = proportion of members under 1818") or use the conventional symbols pp, λ\lambda, μ\mu in the hypotheses. Writing hypotheses in words only, without a parameter, is usually not accepted.
  • State the distribution under H0H_0 explicitly. This line is often worth a mark on its own.
  • Show the comparison as an inequality: "0.0442<0.050.0442 < 0.05" or "−2.0<−1.96-2.0 < -1.96". The examiner needs to see that you compared like with like.
  • Conclusions should mention the significance level and the context, and use non-definite language.
  • If a question says "use the 2.5%2.5\% significance level" for a one-tailed test, the whole 2.5%2.5\% is in one tail. If it says "two-tailed test at 5%5\%", each tail has 2.5%2.5\%.
  • Inconsistent working (for example, the right probability but a decision that contradicts it) loses both the decision and the conclusion marks.
Summary
  • A test assumes H0H_0 is true and asks whether the data would then be surprising.
  • H0H_0 gives one value of a population parameter; H1H_1 is >>, << (one-tailed) or ≠\ne (two-tailed), chosen from the question's wording.
  • The significance level is the probability of rejecting H0H_0 when it is true.
  • Decide by comparing P(result at least as extreme)P(\text{result at least as extreme}) with the level (halved per tail for two-tailed), or by checking whether the test statistic is in the critical region.
  • For discrete distributions, the critical region has probability as close as possible to, but not above, the significance level.
  • Conclude with a decision about H0H_0 and a non-definite statement in context.
  • Rejecting H0H_0 is not proof, and not rejecting it does not show it is true.

Practice questions

Question
  1. For each situation, state the null and alternative hypotheses, defining any parameter used. (a) A seed company claims that 85%85\% of its seeds germinate. A gardener thinks the proportion is lower. (b) The number of accidents on a road averages 2.52.5 per month. A council wants to know whether new speed cameras have changed the rate. (c) A machine fills packets with a mean of 250250 g. An inspector suspects the mean is more than 250250 g.
  2. Explain what is meant by a "5%5\% significance level" in a hypothesis test.
  3. A coin is tossed 1212 times and shows 1010 heads. Test at the 5%5\% level whether the coin is biased towards heads.
  4. A coin is tossed 1212 times and shows 99 heads. Test at the 10%10\% significance level whether the coin is biased.
  5. In a test of H0:p=0.25H_0: p = 0.25 against H1:p>0.25H_1: p > 0.25, a student rejects H0H_0 at the 5%5\% level and writes: "This proves that the proportion has increased to 0.40.4, the proportion in my sample." Give two criticisms of this conclusion.
  6. The masses of bags of rice are normally distributed with standard deviation 66 g. A random sample of 1616 bags has mean mass 47.547.5 g. Test H0:μ=50H_0: \mu = 50 against H1:μ<50H_1: \mu < 50 at (a) the 1%1\% significance level, (b) the 5%5\% significance level.
  7. In a two-tailed test at the 5%5\% significance level, the probability of the observed value or a more extreme value in the same tail, assuming H0H_0, is 0.0320.032. State the conclusion of the test, with a reason.
  8. Faults occur at random in cable at an average rate of 0.80.8 per 100100 m. After a change in production, a 500500 m length of cable is found to contain 11 fault. (a) Test at the 5%5\% significance level whether the fault rate has decreased. (b) Find the critical region for this test.
Answers
  1. (a) pp = proportion of the company's seeds that germinate. H0:p=0.85H_0: p = 0.85, H1:p<0.85H_1: p < 0.85. (b) λ\lambda = mean number of accidents per month. H0:λ=2.5H_0: \lambda = 2.5, H1:λ≠2.5H_1: \lambda \ne 2.5. (c) μ\mu = population mean mass of packets in grams. H0:μ=250H_0: \mu = 250, H1:μ>250H_1: \mu > 250.

  2. It is the probability of rejecting the null hypothesis when it is in fact true (at most 5%5\% for a discrete test statistic). Equivalently, H0H_0 is rejected only if a result as extreme as the one observed would happen less than 5%5\% of the time when H0H_0 is true.

  3. pp = probability of a head. H0:p=0.5H_0: p = 0.5, H1:p>0.5H_1: p > 0.5. Under H0H_0, X∼B(12,0.5)X \sim B(12, 0.5). P(X≥10)=(1210)+(1211)+(1212)212=66+12+14096=0.0193P(X \ge 10) = \dfrac{\binom{12}{10} + \binom{12}{11} + \binom{12}{12}}{2^{12}} = \dfrac{66 + 12 + 1}{4096} = 0.0193. 0.0193<0.050.0193 < 0.05: reject H0H_0. There is evidence at the 5%5\% level that the coin is biased towards heads.

  4. H0:p=0.5H_0: p = 0.5, H1:p≠0.5H_1: p \ne 0.5. Under H0H_0, X∼B(12,0.5)X \sim B(12, 0.5). P(X≥9)=220+66+12+14096=2994096=0.0730P(X \ge 9) = \dfrac{220 + 66 + 12 + 1}{4096} = \dfrac{299}{4096} = 0.0730. Two-tailed at 10%10\%, so compare with 0.050.05: 0.0730>0.050.0730 > 0.05. Do not reject H0H_0. There is insufficient evidence at the 10%10\% level that the coin is biased.

  5. A test never proves anything; the conclusion should say there is evidence that the proportion has increased. Also, the test says nothing about the size of the new proportion: 0.40.4 is a sample value, and the population proportion is not known to equal it.

  6. Under H0H_0, Xˉ∼N(50,3616)\bar{X} \sim N\left(50, \dfrac{36}{16}\right), standard deviation 1.51.5. z=47.5−501.5=−1.667z = \dfrac{47.5 - 50}{1.5} = -1.667. (a) Critical value −2.326-2.326. Since −1.667>−2.326-1.667 > -2.326, do not reject H0H_0: insufficient evidence at the 1%1\% level that the mean mass is less than 5050 g. (b) Critical value −1.645-1.645. Since −1.667<−1.645-1.667 < -1.645, reject H0H_0: evidence at the 5%5\% level that the mean mass is less than 5050 g.

  7. In a two-tailed test at 5%5\%, each tail has 0.0250.025. Since 0.032>0.0250.032 > 0.025, do not reject H0H_0: there is insufficient evidence that the parameter has changed.

  8. (a) λ\lambda = mean number of faults per 500500 m. H0:λ=4H_0: \lambda = 4, H1:λ<4H_1: \lambda < 4. Under H0H_0, X∼Po(4)X \sim \text{Po}(4). P(X≤1)=e−4(1+4)=0.0916>0.05P(X \le 1) = e^{-4}(1 + 4) = 0.0916 > 0.05. Do not reject H0H_0: insufficient evidence at the 5%5\% level that the fault rate has decreased. (b) P(X=0)=e−4=0.0183<0.05P(X = 0) = e^{-4} = 0.0183 < 0.05 and P(X≤1)=0.0916>0.05P(X \le 1) = 0.0916 > 0.05, so the critical region is X=0X = 0.

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action