Hypothesis Tests for a Binomial Proportion
Is a coin biased? Has a new treatment raised the recovery rate? Do fewer customers now use a discount code? Each of these is a claim about a probability , and each can be tested with a single observation: the number of successes in a sample of size . That number has a binomial distribution, so the test uses exact binomial probabilities. Paper 6 asks you to carry out such tests, to find critical regions and actual significance levels, and to handle two-tailed tests, usually as a prelude to finding the probabilities of Type I and Type II errors.
The set-up
A sample of independent trials is observed, each with the same probability of success. The test statistic is the number of successes, . If the null hypothesis is true, then
The alternative hypothesis is , or , depending on the wording of the question (see the nature of hypothesis testing). Large values of are evidence for ; small values are evidence for .
There are no binomial tables in the 9709 formula booklet. You calculate probabilities from
or with your calculator's binomial functions. Questions are usually designed so that the tail you need has only a few terms, but either way you must show the expression you are evaluating.
Method 1: compare a probability with the significance level
Calculate the probability, under , of the observed value or anything more extreme in the direction of .
- For with observed value , find .
- For with observed value , find .
- For , find the tail on the side of the observed value (compare with the mean to decide which), and compare with half the significance level.
If the probability is less than the significance level, reject .
The reason for using rather than is that any single value can be unlikely. With , even the most likely outcome has a small probability. What matters is whether the observation is far out in the tail, and a tail probability measures exactly that.
- Define in context and write and .
- State "under , ", defining .
- Write down the tail probability for the observed value, in the direction of , as a sum of terms or as .
- Compare with the significance level (half of it for a two-tailed test), showing the inequality.
- State whether is rejected, then conclude in context without certainty.
In the general population of people are left-handed. A researcher suspects that left-handedness is less common among basketball players. In a random sample of basketball players, are left-handed. Test the researcher's suspicion at the significance level.
Solution
Let be the proportion of basketball players who are left-handed.
, .
Let be the number of left-handed players in the sample. Under , .
, so do not reject .
There is insufficient evidence at the level that left-handedness is less common among basketball players.
A standard treatment for a condition is successful for of patients. A new treatment is tried on a random sample of patients and is successful for of them. Test at the significance level whether the new treatment has a higher success rate.
Solution
Let be the probability that the new treatment is successful for a patient.
, .
Let be the number of successes. Under , .
, so reject .
There is evidence at the level that the new treatment has a higher success rate than the standard treatment.
Method 2: find the critical region
The critical region is the set of values of that would lead to rejecting . Because is discrete, you cannot usually hit the significance level exactly. The rule is:
- Lower tail, : the critical region is , where is the largest value with .
- Upper tail, : the critical region is , where is the smallest value with .
- Two-tailed, : find a lower region and an upper region, each with probability as close as possible to, but not more than, .
The actual significance level is the probability of the critical region when is true. It is at most , and usually less.
To justify a critical region you must show the cumulative probability for the boundary value and for the next value along, which is too big. That pair of numbers is what proves is right.
The graph shows as bars of width , with the lower-tail critical region at the level shaded. Its total probability is ; including would take it to , which is too much.
- State , and the distribution under .
- Build up the cumulative probabilities from the end of the relevant tail: , , , ... (or , , ... for the upper tail).
- Stop when the cumulative probability first exceeds the significance level (or half of it, for each tail of a two-tailed test).
- State the critical region, quoting the last probability that was within the limit and the first that was not.
- If asked, the actual significance level is the probability of the critical region.
It is claimed that of customers at a shop use a discount code. The manager suspects that the proportion is lower. She will test the claim at the significance level using a random sample of customers.
(a) Find the critical region for the test.
(b) State the actual significance level.
(c) In the sample, customers use a discount code. Carry out the test.
Solution
(a) Let be the proportion of customers who use a code. , . Under , .
but , so the critical region is .
(b) The actual significance level is , that is .
(c) is not in the critical region, so do not reject . There is insufficient evidence at the level that fewer than of customers use a discount code.
A survey says of households in a town own a cat. A vet wants to test whether this proportion has changed, using a random sample of households and a significance level.
(a) Find the critical region.
(b) Find the actual significance level of the test.
Solution
(a) , . Under , . Each tail may have probability at most .
Lower tail:
So the lower part of the critical region is .
Upper tail:
So the upper part is .
The critical region is or .
(b) The actual significance level is
The two tails need not be equal, and the actual level is below because neither tail can reach exactly .
A company claims that of customers prefer its brand. A rival believes the proportion is lower and plans a test at the significance level, based on the number of customers in a random sample of who prefer the brand.
(a) Show that if , the test can never lead to rejecting the company's claim.
(b) Find the smallest value of for which it is possible to reject the claim.
(c) The rival uses . Find the critical region and the actual significance level.
(d) In fact, of the customers prefer the brand. State the conclusion of the test.
Solution
(a) , . The most extreme result possible is . With ,
so even the most extreme result is not in the critical region, and can never be rejected.
(b) Rejection is possible only if .
Check: and . The smallest value is .
(c) Under , .
The critical region is , and the actual significance level is , that is .
Notice how far below this is: is only just over , but it is over, so cannot be included.
(d) is not in the critical region. Do not reject : there is insufficient evidence at the level that fewer than of customers prefer the brand.
- Using instead of a tail probability. For "increased", find ; for "decreased", .
- Off-by-one errors in the complement. , not .
- Choosing the wrong tail in a two-tailed test. Compare the observed value with the mean . Above the mean, use the upper tail; below it, the lower tail.
- Not halving the level for a two-tailed test. Each tail of a two-tailed test has .
- Picking the critical value whose probability is closest to . The critical region's probability must not exceed the significance level, even if a larger region would be closer.
- Justifying a critical region with one probability. Show the boundary value's cumulative probability () and the next one ().
- Saying the actual significance level is . For a discrete test it is the probability of the critical region, such as .
- Define in context ("the proportion of customers who use a code"). Hypotheses in terms of or the sample are not accepted.
- "Under , " earns a mark. Write it.
- Write tail probabilities as expressions before evaluating. If you use a calculator's cumulative function, still state which probability you found, such as .
- For "find the critical region", the answer is a set of values of , written as an inequality such as or . Do not give a probability as the answer.
- The comparison and conclusion follow the same rules as every test: show the inequality, state the decision about , and conclude in context with "evidence" language.
- Many questions continue by asking for , which is the actual significance level you have just found.
- Test statistic: the number of successes ; under , .
- Compare or (in the direction of ) with the significance level, halved for each tail of a two-tailed test.
- Lower-tail critical region: largest with . Upper-tail: smallest with .
- Show the cumulative probabilities either side of the boundary.
- The actual significance level is the probability of the critical region under , at most .
- If even the most extreme outcome has probability above , the test can never reject .
- Conclude in context with non-definite language.
Practice questions
- Find the critical region for a test of against at the significance level, using . State the actual significance level.
- A blood group is found in of the population. In a random sample of people from a particular region, have this blood group. Test at the significance level whether the blood group is less common in this region.
- A spinner is designed so that the probability of red is . In spins, red occurs times. Test at the significance level whether the spinner is biased towards red.
- A coin is to be tested for bias using tosses and a two-tailed test at the level. Find the critical region and the actual significance level.
- A test of against at the significance level is based on a random sample of size . Find the smallest value of for which could be rejected.
- It is thought that of emails received by a company are spam. After a filter is updated, a manager believes the proportion of spam getting through has increased. A random sample of emails is checked, and a test is carried out at the significance level. (a) Find the critical region. (b) The sample contains spam emails. State the conclusion of the test.
- A politician claims that at least of voters support her. An opponent believes that support is lower. In a random sample of voters, support the politician. Test the opponent's belief at the significance level.
- A test of against uses a random sample of size , and is rejected if or . (a) Find the significance level of the test. (b) Show that this is the critical region for a two-tailed test at the significance level. (c) In the sample, . State the conclusion.
Answers
-
Under , . and . Critical region ; actual significance level .
-
= proportion in the region with the blood group. , . Under , . . : reject . There is evidence at the level that the blood group is less common in this region.
-
= probability of red. , . Under , . . : reject . There is evidence at the level that the spinner is biased towards red.
-
, , , each tail at most . ; . Lower region ; by symmetry the upper region is . Critical region or ; actual significance level .
-
Need , so . Check: , . Smallest .
-
(a) , , . ; . Critical region . (b) is not in the critical region. Do not reject : insufficient evidence at the level that the proportion of spam getting through has increased.
-
= proportion of voters who support her. , . Under , . (sum of the six terms to ). : do not reject . There is insufficient evidence at the level that her support is lower than .
-
(a) Under , . and . Significance level . (b) Lower: and . Upper: and . So these are the largest regions in each tail with probability at most . (c) is in the critical region. Reject : there is evidence at the level that the proportion is not . (The result is in the upper tail, which suggests that if the proportion has changed, it has increased.)