Populations and Samples
Almost everything in the second half of Paper 6 rests on one idea: you want to know something about a whole population, but you can only measure a sample. Estimates, confidence intervals and hypothesis tests all turn sample data into a statement about the population, and they only work if the sample is chosen at random. This note covers the vocabulary, why randomness matters, how random numbers produce a random sample, and how to explain in context why a sampling method is unsatisfactory, which is a short but regular question on the paper.
Populations, samples and the words around them
Start with a concrete case. A factory makes light bulbs a week, and the manager wants to know their mean lifetime. The bulbs are the population. Testing a bulb means running it until it fails, so the manager tests of them. Those are the sample.
- The population is the whole set of items, people or values that you want to know about.
- A census collects data from every member of the population.
- A sample is a subset of the population from which data are actually collected.
- A sampling frame is a list or map of the members of the population, from which the sample is chosen.
- A parameter is a numerical property of the population, such as the population mean , variance or proportion . It is a fixed number, usually unknown.
- A statistic is a numerical value calculated from a sample, such as the sample mean . It varies from sample to sample.
The distinction between a parameter and a statistic is the heart of the topic. The population mean lifetime of the bulbs is one fixed number that nobody knows. The mean of the tested bulbs is known, but a different bulbs would have given a different value. Statistics is the business of using the second to say something sensible about the first.
Why take a sample at all?
A census gives exact information about the population, so why not always use one? Because in practice a census is often impossible or unwise.
| Reason to sample | Example |
|---|---|
| Testing destroys the item | Lifetimes of bulbs, breaking strength of cables, taste-testing a batch of food |
| The population is very large or infinite | All possible throws of a die, every fish in a lake |
| A census would take too long | Opinion poll needed before an election next week |
| A census would cost too much | Measuring the height of every adult in a country |
| The population is not fully listable | All future customers of a shop |
The price of sampling is uncertainty. A sample statistic will not usually equal the population parameter exactly, and a different sample would give a different value. The later notes in this unit measure that uncertainty precisely.
A census is sensible when the population is small and every member can be measured without harm, for example the exam marks of the students in one class.
Why the sample must be random
A sample is only useful if it represents the population. The surest way to get a representative sample, and the only way that lets you calculate how reliable your conclusions are, is to choose it at random.
A random sample of size is a sample chosen so that every possible set of members of the population has the same chance of being selected. In particular, every member of the population has the same chance of being chosen.
A sampling method is biased if it tends to favour some members of the population over others, so that the sample systematically misrepresents the population.
Bias is not the same as bad luck. A random sample might by chance contain unusually long-lasting bulbs, but if you repeated the sampling many times, the errors would average out. A biased method makes the same kind of error every time. Interviewing shoppers in an expensive department store will over-estimate the average income of a town on every occasion, however many people you ask. A bigger sample does not cure bias; it only makes you more confident in the wrong answer.
The second reason randomness matters is mathematical. The results in the following notes, such as and , the Central Limit Theorem, confidence intervals and hypothesis tests, all assume the sample values are independent observations of the same random variable. A random sample delivers exactly that. A non-random sample breaks every one of those calculations.
The syllabus says that knowledge of particular named sampling methods, such as quota or stratified sampling, is not required. You need an elementary understanding of how random numbers produce a random sample, and the ability to explain in simple terms why a given method may be unsatisfactory. Named methods may appear in a question's wording, but you will not be asked to define or carry them out.
Using random numbers to choose a sample
Random numbers are digits in which each of is equally likely and independent of the others. They come from a calculator (the Ran# or RanInt function), a computer, or a printed table.
- Obtain a sampling frame: a list of the whole population.
- Number the members, using the same number of digits for each: to , or to .
- Generate random numbers with that many digits (two-digit numbers for to ).
- Ignore any number outside the range, and ignore repeats (so no member is chosen twice).
- Continue until different members have been chosen. These form the sample.
Ignoring out-of-range numbers is what keeps the method fair. If a list has members and you read two-digit numbers, then and to simply do not correspond to anyone, and skipping them leaves every member with the same chance.
A club has members, numbered to . A sample of members is to be chosen using the following random digits, read from left to right in pairs.
Find the numbers of the members chosen.
Solution
Split into pairs: .
- : out of range, ignore.
- : choose member .
- : choose member .
- : out of range, ignore.
- : repeat, ignore.
- : choose member .
- : repeat, ignore.
- : choose member .
The sample is members , , and .
A school has students. The head teacher wants a random sample of students to complete a questionnaire. Describe how random numbers can be used to choose the sample.
Solution
Use the school register as the sampling frame and number the students from to . Use a calculator or computer to generate three-digit random numbers. Ignore , any number greater than , and any number that has already appeared. Continue until different numbers have been obtained, and choose the students with those numbers.
Full marks for a description like this need three things: a list of the population, each member numbered; random numbers generated; and how out-of-range numbers and repeats are handled, or equivalently that the process continues until different members are chosen.
Explaining why a method is unsatisfactory
The most common sampling question gives a method in context and asks you to explain why it is unsatisfactory, or why it does not give a random sample. The answer is almost always that some members of the population cannot be chosen, or are more likely to be chosen than others, and that this matters because those members are likely to differ in the very thing being measured.
- Identify the population the question is really about.
- Ask who in that population cannot be selected by this method.
- Ask who is more likely to be selected, or who chooses whether to take part.
- Say how the excluded or favoured group is likely to differ in the quantity being measured, so the results will be too high or too low.
- Write all of this in the context of the question.
The usual sources of bias are worth recognising on sight.
- Convenience: choosing whoever is nearest or easiest to reach (the first people through the door, the plants at the edge of the field).
- Time and place: surveying at one location or one time excludes everyone who is elsewhere then.
- Self-selection: when people choose whether to respond (phone-ins, online polls, a questionnaire left on a desk), those with strong opinions are over-represented.
- An incomplete sampling frame: a list that leaves out part of the population, such as a phone directory that excludes people with only mobile phones.
- Hidden patterns: taking every 10th item from a production line can coincide with a machine fault that recurs every 10 items.
- Non-response: if many people chosen do not reply, the replies may come from an unrepresentative group.
A researcher wants to estimate the mean time that working adults in a town spend travelling to work. She interviews people at the town's railway station between am and am on a Monday.
(a) Give two reasons why this sample is unlikely to be representative.
(b) Suggest a better method.
Solution
(a) Only people who travel by train can be chosen, so people who drive, cycle, walk or work from home are excluded. Train journeys are likely to be longer than many of these, so the estimate is likely to be too high.
Only people travelling between am and am on a Monday can be chosen, so people with different working hours, or who do not work on Mondays, are excluded.
(b) Obtain a list of all working adults in the town (for example from a register of residents), number them, and use random numbers to choose of them to ask.
A council wants to estimate the proportion of households in a town that would use a new recycling service. Three methods are suggested.
A: Phone numbers chosen at random from the local landline directory.
B: Ask every household in one street, chosen at random.
C: Put a notice in the local newspaper asking households to reply online.
For each method, give one reason why it may not give a representative sample.
Solution
A: Households without a landline, or whose number is not listed, cannot be chosen. These may be younger or less settled households, whose views on recycling may differ.
B: Households in one street tend to be similar (similar housing, income, and the same existing collection arrangements), so their views will not represent the whole town. Every household in the town does have a chance of selection, but not every set of households: only sets that make up a single street can occur.
C: The sample is self-selected. Only households that read the paper can respond, and of those, the households most interested in recycling are the most likely to reply, so the proportion is likely to be over-estimated.
Every member equally likely is not enough
The definition of a random sample asks for more than "every member has an equal chance". Every possible sample of size must be equally likely. Some methods pass the first test and fail the second.
A school has students listed alphabetically in a register, students to a page over pages. To choose a sample of students, a teacher picks one page at random and takes all students on that page.
(a) Show that every student has the same probability of being chosen.
(b) Explain why this is nevertheless not a random sample of size .
(c) Explain why the sample might be biased in a survey about family background.
Solution
(a) A student is chosen exactly when their page is chosen, which happens with probability . This is the same for every student, and , the same as in a genuinely random sample.
(b) Only different samples are possible: the pages. A sample containing, say, the first student on page and the last student on page can never occur. In a random sample every set of students would be possible, so this method is not random.
(c) Students on one page have surnames beginning with the same few letters. Students who share a surname may be siblings, and surnames are linked to family and cultural background, so the sample could over-represent particular families or backgrounds.
A statistic is a random variable
Take two random samples of bulbs and you will get two different sample means. Before the sample is taken, the sample mean is uncertain: it is a random variable, written with a capital letter, . Once a particular sample has been measured, its mean is a number, written .
Because is a random variable, it has a distribution, with its own mean and variance. That distribution, the sampling distribution of the mean, tells you how far a sample mean is likely to be from the population mean, and is the subject of the distribution of the sample mean. The same idea gives unbiased estimates of and .
- Writing "it is not random" and stopping. That earns nothing. Name who is excluded or over-represented, in context, and say how this might affect the result.
- Thinking a larger sample removes bias. A biased method gives a biased result whatever the sample size. Size reduces random variation, not bias.
- Confusing "random" with "haphazard". Picking items "without thinking" or "the ones that look typical" is not random. Random selection uses random numbers or an equivalent physical process.
- Forgetting repeats and out-of-range numbers. A description of random number sampling must say what to do with them.
- Confusing a parameter with a statistic. is a fixed (unknown) population value; is calculated from a sample and changes from sample to sample.
- Assuming equal chances for individuals is enough. A random sample needs every possible sample to be equally likely.
- "Explain why the method is unsatisfactory" is usually 1 or 2 marks. One clear, contextual reason per mark: "only people who travel by train can be chosen, so people who drive are excluded".
- "Describe how to use random numbers" needs: number the population (from a list), generate random numbers, ignore numbers out of range and repeats, continue until members are chosen.
- "Why is a sample used rather than a census?" Give a reason that fits the context: testing is destructive, or the population is too large.
- "State what is meant by a population parameter" or "a random sample": use the definitions above.
- In later parts of a question, the requirement that "the sample is random" is often the assumption you are asked to state, for example before using the Central Limit Theorem or a hypothesis test.
- Population: everything of interest. Sample: the part actually measured. Census: measure everything.
- Parameters (, , ) describe the population and are fixed; statistics (, ) come from samples and vary.
- Sample when testing is destructive, or the population is too large, slow or expensive to measure.
- A random sample makes every possible sample of size equally likely; it avoids bias and justifies the later calculations.
- Random numbers: number the sampling frame, generate numbers, ignore out-of-range values and repeats, continue until are chosen.
- Criticise a method by naming who cannot be chosen or is more likely to be chosen, in context, and how that distorts the result.
- A sample mean is a random variable with its own distribution.
Practice questions
- A manufacturer of matches wants to know the proportion of its matches that light first time. Explain why a sample must be used rather than a census.
- State the difference between a parameter and a statistic, using the mean mass of apples from an orchard as your example.
- A firm has employees. Describe how to use random numbers to choose a sample of employees.
- A list of items is numbered to . Use the random digits , read in pairs from left to right, to choose items.
- To estimate the mean number of hours of television watched per week by children in a city, a researcher asks the children at one primary school. Give two reasons why the sample may not be representative.
- A radio station asks listeners to phone in to say whether they support a new road. Of callers, oppose it. Explain why it would be unwise to conclude that of the region's population oppose the road.
- A quality inspector tests every 25th bottle filled by a machine that has filling nozzles used in strict rotation. Explain why this method could fail to detect a faulty nozzle, or exaggerate its effect.
- A college has students. Method P: put all names in a box and draw without replacement. Method Q: number the students to , choose one of the first numbers at random, and then take every 20th student after it. (a) Show that under each method, each student has probability of being chosen. (b) Explain why only method P gives a random sample, and give a situation in which method Q could produce a biased sample.
Answers
-
Testing whether a match lights destroys it, so a census would leave nothing to sell. (The population is also very large.)
-
The population mean mass of all the apples in the orchard is a parameter: a single fixed value, unknown in practice. The mean mass of a sample of, say, apples is a statistic: it is calculated from the sample and would be different for a different sample of .
-
Obtain a list of all employees and number them to . Generate three-digit random numbers using a calculator or computer. Ignore , numbers above and repeats. Continue until different numbers are obtained and choose the corresponding employees.
-
Pairs: . Ignore (out of range) and the second (repeat). The items chosen are , , and .
-
Only children at that school can be chosen; children at other schools, which may be in different areas with different family backgrounds, are excluded. Also only primary-age children are included, so older children (who may watch different amounts) are excluded.
-
The callers are self-selected: only people who listen to that station can call, and people with strong opinions, especially opponents, are much more likely to phone. The sample is biased, so is likely to over-estimate the proportion who oppose the road.
-
Since is a multiple of , every 25th bottle comes from the same nozzle. If that nozzle works, faults in the other four are never seen; if that nozzle is the faulty one, every tested bottle is faulty and the fault rate is greatly exaggerated.
-
(a) P: by symmetry each student is equally likely, and of are chosen, so the probability is . Q: a student is chosen exactly when the starting number is the one in their position within each block of , which has probability . (b) In P every set of students is equally likely. In Q only different samples are possible (one for each starting number); for example two consecutively numbered students can never both be chosen, so Q is not a random sample. If the list were arranged in tutor groups of with, say, the group representative listed first, then Q would either pick all the representatives or none of them.