The Nature of Hypothesis Testing
A hypothesis test is a formal way of deciding whether data give real evidence against a claim, or whether what you observed could easily have happened by chance. A manufacturer claims that only of its components are faulty; a sample has faulty out of . Is the claim false, or was this an unlucky sample? Every S2 test answers a question like this with the same logic and the same written structure, and every Paper 6 has at least one full test, usually worth or marks. This note sets out the language and the logic; the notes that follow apply them to binomial, Poisson and normal tests.
The logic of a test
Think of a courtroom. The defendant is presumed innocent, and is convicted only if the evidence would be very unlikely if they were innocent. A hypothesis test works the same way.
- Start by assuming the claim is true. This is the null hypothesis.
- Ask how likely a result at least as extreme as the one observed would be, if the claim were true.
- If that probability is very small (below a threshold fixed in advance), the data are hard to explain under the claim, so reject it. Otherwise, the data are consistent with the claim and you do not reject it.
Here is the whole idea in one small example. A coin is claimed to be fair, but you suspect it favours heads. You toss it times and get heads. If the coin really is fair, the number of heads satisfies , and
Nine or more heads would happen only about of the time with a fair coin. That is strong evidence that the coin is biased towards heads. It is not proof: fair coins do sometimes give heads in tosses. A test controls how often this kind of wrong conclusion happens, but cannot rule it out.
The vocabulary
The syllabus lists the terms you must understand and use. Learn them precisely; definitions are sometimes asked for directly.
- The null hypothesis, , is the claim being tested, assumed true until the evidence says otherwise. It always states a single value of a population parameter, such as , or .
- The alternative hypothesis, , says how the parameter differs from that value if is false: , or .
- A one-tailed test has an alternative hypothesis in one direction ( or ). A two-tailed test has with , so extreme results in either direction count as evidence.
- The test statistic is the quantity calculated from the sample and used to decide: the number of successes, the number of events, or the sample mean (often converted to a -value).
- The significance level is the threshold probability for rejecting , such as or . It is the probability of rejecting when is true (for a discrete test statistic, the largest such probability allowed).
- The critical region (or rejection region) is the set of values of the test statistic for which is rejected. The acceptance region is the set of values for which is not rejected. The critical value is the boundary value of the critical region.
Writing the hypotheses
Hypotheses are always about the population parameter, never about the sample. "" is about the proportion of the whole population; the sample result, such as out of , is the evidence.
The question's wording tells you which alternative to use.
| Wording in the question | Alternative hypothesis | Tails |
|---|---|---|
| "has increased", "is more than", "is biased towards", "underestimates" | One (upper) | |
| "has decreased", "is less than", "has improved" (when lower is better) | One (lower) | |
| "has changed", "is different", "is biased" (no direction) | Two |
The direction must be decided by the question, before looking at the data. Choosing a one-tailed test because the sample happens to be high is not allowed.
Two ways to reach a decision
Once , and the distribution of the test statistic under are set out, there are two equivalent ways to decide.
Compare a probability with the significance level. Calculate the probability, assuming , of a result at least as extreme as the one observed, in the direction of . If it is less than the significance level, reject . (This probability is often called the -value, though Paper 6 does not require the term.) For a two-tailed test, compare the probability in the observed tail with half the significance level.
Find the critical region. Work out which values of the test statistic would lead to rejection, then see whether the observed value is among them. This is essential when a question asks for the critical region, or for the probability of a Type I or Type II error.
For a test based on a normal distribution, the critical region is in one or both tails of the curve. The first graph shows a two-tailed test at the level, with in each tail beyond ; the second shows a one-tailed (upper) test at the level, with the whole beyond .
Discrete test statistics
For a binomial or Poisson test statistic, probabilities jump in steps, so a critical region usually cannot have probability exactly . The critical region is chosen to have probability as close as possible to, but not more than, the significance level. Its actual probability is called the actual significance level of the test. You will see this in detail in tests for a binomial proportion.
Writing the conclusion
The conclusion has two parts: the decision about , and what it means in the context of the question. Both are needed for the final mark.
- If the result is in the critical region: "Reject . There is evidence at the level that the coin is biased towards heads."
- If not: "Do not reject . There is insufficient evidence at the level that the coin is biased towards heads." Cambridge mark schemes also accept "accept ", provided the contextual statement is not definite.
Never claim certainty. A test does not prove that is false, and failing to reject does not show that it is true; it only means the data are consistent with it. Avoid words like "proves", "shows that" and "the coin is fair".
Which test, which distribution
| Situation | Test statistic | Distribution under | Note |
|---|---|---|---|
| Proportion, single observation of a count of successes | Number of successes | Binomial proportion | |
| Rate of random events | Number of events | Poisson mean | |
| Large or large | , with continuity correction | Normal approximation | Normal approximations |
| Mean of a population | Sample mean | Population mean |
- Define the parameter in words and state and in symbols: " = probability of a head; , ".
- State the test statistic and its distribution assuming is true: "".
- Calculate the probability of a result at least as extreme as the one observed (or find the critical region).
- Compare explicitly with the significance level (or say whether the observed value is in the critical region): "".
- State the decision about .
- Write the conclusion in context, without certainty.
A coin is tossed times and shows heads. Test at the significance level whether the coin is biased towards heads.
Solution
Let be the probability that the coin shows a head.
, .
Let be the number of heads in tosses. Under , .
, so reject .
There is evidence at the significance level that the coin is biased towards heads.
For each situation, define a parameter, and state suitable null and alternative hypotheses. Say whether the test is one-tailed or two-tailed.
(a) A website claims that of visitors buy something. A manager believes the true figure is lower.
(b) Emails arrive at an average rate of per hour. After a new spam filter is installed, an analyst wants to know whether the rate has changed.
(c) The mean mass of bags of flour is supposed to be g. A trading standards officer suspects that the bags are underweight on average.
Solution
(a) = proportion of all visitors who buy something. , . One-tailed.
(b) = mean number of emails per hour. , . Two-tailed, because "changed" has no direction.
(c) = population mean mass of bags, in grams. , . One-tailed.
In each case gives a single value, so the distribution of the test statistic can be written down exactly.
A die is thrown times and a six appears only once. Test at the significance level whether the die is biased.
Solution
Let be the probability of a six. , (two-tailed).
Let be the number of sixes in throws. Under , . The observed value is below the expected value , so look at the lower tail.
The test is two-tailed, so compare with : . Do not reject .
There is insufficient evidence at the level that the die is biased.
Notice that a one-tailed test of would have rejected , since . The choice of tails matters, which is why it must be made from the question, not from the data.
The volumes of drink in bottles are normally distributed with standard deviation ml. The mean is supposed to be ml. A random sample of bottles has mean volume ml. Test at the significance level whether the mean volume has changed.
Solution
Let be the population mean volume in ml. , (two-tailed).
Under , , standard deviation .
Method 1, comparing with the critical value.
The critical values for a two-tailed test at are . Since , the result is in the critical region.
Method 2, comparing a probability with the level.
Either way, reject . There is evidence at the level that the mean volume of drink in the bottles has changed.
It is known that of the members of a sports club are under . After a recruitment drive, a random sample of members contains who are under .
(a) Test at the significance level whether the proportion of members under has increased.
(b) Find the smallest significance level, correct to 1 decimal place, at which the test in (a) would lead to rejecting .
(c) A colleague instead tests whether the proportion has changed, at the level. State, with a reason, the conclusion of this test.
Solution
(a) Let be the proportion of members who are under .
, .
Under , , where is the number under in the sample.
(The sum has twelve terms; in practice you would use a calculator's binomial cumulative function, and the method mark is for the expression with .)
, so reject . There is evidence at the level that the proportion of members under has increased.
(b) is rejected whenever the significance level exceeds , so the smallest level is .
(c) For a two-tailed test, the upper tail probability is compared with . Since , is not rejected: there is insufficient evidence at the level that the proportion has changed.
- Hypotheses about the sample. "" or "" (the sample proportion) is wrong. Hypotheses are about the population parameter, with the value from the claim.
- Inequalities in . is always an equality. "" is not accepted.
- Comparing the wrong probability. Use the probability of the observed value or more extreme, never on its own. For "increased", that is ; for "decreased", .
- Forgetting to halve the level in a two-tailed test. At the level, each tail gets .
- Definite conclusions. "The coin is biased" or "this proves the mean has changed" loses the mark. Use "there is evidence that" or "there is insufficient evidence that".
- No context. "Reject " alone does not earn the final mark; say what it means in the situation.
- Choosing the tail after seeing the data. The direction of comes from the question's wording.
- Define the parameter in words (" = proportion of members under ") or use the conventional symbols , , in the hypotheses. Writing hypotheses in words only, without a parameter, is usually not accepted.
- State the distribution under explicitly. This line is often worth a mark on its own.
- Show the comparison as an inequality: "" or "". The examiner needs to see that you compared like with like.
- Conclusions should mention the significance level and the context, and use non-definite language.
- If a question says "use the significance level" for a one-tailed test, the whole is in one tail. If it says "two-tailed test at ", each tail has .
- Inconsistent working (for example, the right probability but a decision that contradicts it) loses both the decision and the conclusion marks.
- A test assumes is true and asks whether the data would then be surprising.
- gives one value of a population parameter; is , (one-tailed) or (two-tailed), chosen from the question's wording.
- The significance level is the probability of rejecting when it is true.
- Decide by comparing with the level (halved per tail for two-tailed), or by checking whether the test statistic is in the critical region.
- For discrete distributions, the critical region has probability as close as possible to, but not above, the significance level.
- Conclude with a decision about and a non-definite statement in context.
- Rejecting is not proof, and not rejecting it does not show it is true.
Practice questions
- For each situation, state the null and alternative hypotheses, defining any parameter used. (a) A seed company claims that of its seeds germinate. A gardener thinks the proportion is lower. (b) The number of accidents on a road averages per month. A council wants to know whether new speed cameras have changed the rate. (c) A machine fills packets with a mean of g. An inspector suspects the mean is more than g.
- Explain what is meant by a " significance level" in a hypothesis test.
- A coin is tossed times and shows heads. Test at the level whether the coin is biased towards heads.
- A coin is tossed times and shows heads. Test at the significance level whether the coin is biased.
- In a test of against , a student rejects at the level and writes: "This proves that the proportion has increased to , the proportion in my sample." Give two criticisms of this conclusion.
- The masses of bags of rice are normally distributed with standard deviation g. A random sample of bags has mean mass g. Test against at (a) the significance level, (b) the significance level.
- In a two-tailed test at the significance level, the probability of the observed value or a more extreme value in the same tail, assuming , is . State the conclusion of the test, with a reason.
- Faults occur at random in cable at an average rate of per m. After a change in production, a m length of cable is found to contain fault. (a) Test at the significance level whether the fault rate has decreased. (b) Find the critical region for this test.
Answers
-
(a) = proportion of the company's seeds that germinate. , . (b) = mean number of accidents per month. , . (c) = population mean mass of packets in grams. , .
-
It is the probability of rejecting the null hypothesis when it is in fact true (at most for a discrete test statistic). Equivalently, is rejected only if a result as extreme as the one observed would happen less than of the time when is true.
-
= probability of a head. , . Under , . . : reject . There is evidence at the level that the coin is biased towards heads.
-
, . Under , . . Two-tailed at , so compare with : . Do not reject . There is insufficient evidence at the level that the coin is biased.
-
A test never proves anything; the conclusion should say there is evidence that the proportion has increased. Also, the test says nothing about the size of the new proportion: is a sample value, and the population proportion is not known to equal it.
-
Under , , standard deviation . . (a) Critical value . Since , do not reject : insufficient evidence at the level that the mean mass is less than g. (b) Critical value . Since , reject : evidence at the level that the mean mass is less than g.
-
In a two-tailed test at , each tail has . Since , do not reject : there is insufficient evidence that the parameter has changed.
-
(a) = mean number of faults per m. , . Under , . . Do not reject : insufficient evidence at the level that the fault rate has decreased. (b) and , so the critical region is .