Populations and Samples

A2 · S2 · 18 min

Almost everything in the second half of Paper 6 rests on one idea: you want to know something about a whole population, but you can only measure a sample. Estimates, confidence intervals and hypothesis tests all turn sample data into a statement about the population, and they only work if the sample is chosen at random. This note covers the vocabulary, why randomness matters, how random numbers produce a random sample, and how to explain in context why a sampling method is unsatisfactory, which is a short but regular question on the paper.

Populations, samples and the words around them

Start with a concrete case. A factory makes 20 00020\,000 light bulbs a week, and the manager wants to know their mean lifetime. The bulbs are the population. Testing a bulb means running it until it fails, so the manager tests 5050 of them. Those 5050 are the sample.

Definition
  • The population is the whole set of items, people or values that you want to know about.
  • A census collects data from every member of the population.
  • A sample is a subset of the population from which data are actually collected.
  • A sampling frame is a list or map of the members of the population, from which the sample is chosen.
  • A parameter is a numerical property of the population, such as the population mean μ\mu, variance σ2\sigma^2 or proportion pp. It is a fixed number, usually unknown.
  • A statistic is a numerical value calculated from a sample, such as the sample mean xˉ\bar{x}. It varies from sample to sample.

The distinction between a parameter and a statistic is the heart of the topic. The population mean lifetime of the bulbs is one fixed number that nobody knows. The mean of the 5050 tested bulbs is known, but a different 5050 bulbs would have given a different value. Statistics is the business of using the second to say something sensible about the first.

Why take a sample at all?

A census gives exact information about the population, so why not always use one? Because in practice a census is often impossible or unwise.

Reason to sampleExample
Testing destroys the itemLifetimes of bulbs, breaking strength of cables, taste-testing a batch of food
The population is very large or infiniteAll possible throws of a die, every fish in a lake
A census would take too longOpinion poll needed before an election next week
A census would cost too muchMeasuring the height of every adult in a country
The population is not fully listableAll future customers of a shop

The price of sampling is uncertainty. A sample statistic will not usually equal the population parameter exactly, and a different sample would give a different value. The later notes in this unit measure that uncertainty precisely.

A census is sensible when the population is small and every member can be measured without harm, for example the exam marks of the 2424 students in one class.

Why the sample must be random

A sample is only useful if it represents the population. The surest way to get a representative sample, and the only way that lets you calculate how reliable your conclusions are, is to choose it at random.

Definition

A random sample of size nn is a sample chosen so that every possible set of nn members of the population has the same chance of being selected. In particular, every member of the population has the same chance of being chosen.

A sampling method is biased if it tends to favour some members of the population over others, so that the sample systematically misrepresents the population.

Bias is not the same as bad luck. A random sample might by chance contain unusually long-lasting bulbs, but if you repeated the sampling many times, the errors would average out. A biased method makes the same kind of error every time. Interviewing shoppers in an expensive department store will over-estimate the average income of a town on every occasion, however many people you ask. A bigger sample does not cure bias; it only makes you more confident in the wrong answer.

The second reason randomness matters is mathematical. The results in the following notes, such as E(Xˉ)=μE(\bar{X}) = \mu and Var(Xˉ)=σ2n\text{Var}(\bar{X}) = \dfrac{\sigma^2}{n}, the Central Limit Theorem, confidence intervals and hypothesis tests, all assume the sample values are independent observations of the same random variable. A random sample delivers exactly that. A non-random sample breaks every one of those calculations.

What the syllabus does and does not require

The syllabus says that knowledge of particular named sampling methods, such as quota or stratified sampling, is not required. You need an elementary understanding of how random numbers produce a random sample, and the ability to explain in simple terms why a given method may be unsatisfactory. Named methods may appear in a question's wording, but you will not be asked to define or carry them out.

Using random numbers to choose a sample

Random numbers are digits in which each of 0,1,…,90, 1, \dots, 9 is equally likely and independent of the others. They come from a calculator (the Ran# or RanInt function), a computer, or a printed table.

Choosing a random sample with random numbers
  1. Obtain a sampling frame: a list of the whole population.
  2. Number the members, using the same number of digits for each: 0101 to 8080, or 001001 to 640640.
  3. Generate random numbers with that many digits (two-digit numbers for 0101 to 8080).
  4. Ignore any number outside the range, and ignore repeats (so no member is chosen twice).
  5. Continue until nn different members have been chosen. These form the sample.

Ignoring out-of-range numbers is what keeps the method fair. If a list has 8080 members and you read two-digit numbers, then 0000 and 8181 to 9999 simply do not correspond to anyone, and skipping them leaves every member with the same chance.

Reading a random number table

A club has 8080 members, numbered 0101 to 8080. A sample of 44 members is to be chosen using the following random digits, read from left to right in pairs.

8 3 0 74 1 9 50 7 2 64 1 3 88\,3\,0\,7\quad 4\,1\,9\,5\quad 0\,7\,2\,6\quad 4\,1\,3\,8

Find the numbers of the members chosen.

Solution

Split into pairs: 83, 07, 41, 95, 07, 26, 41, 3883,\ 07,\ 41,\ 95,\ 07,\ 26,\ 41,\ 38.

  • 8383: out of range, ignore.
  • 0707: choose member 0707.
  • 4141: choose member 4141.
  • 9595: out of range, ignore.
  • 0707: repeat, ignore.
  • 2626: choose member 2626.
  • 4141: repeat, ignore.
  • 3838: choose member 3838.

The sample is members 0707, 4141, 2626 and 3838.

Describing the method in words

A school has 640640 students. The head teacher wants a random sample of 3030 students to complete a questionnaire. Describe how random numbers can be used to choose the sample.

Solution

Use the school register as the sampling frame and number the students from 001001 to 640640. Use a calculator or computer to generate three-digit random numbers. Ignore 000000, any number greater than 640640, and any number that has already appeared. Continue until 3030 different numbers have been obtained, and choose the students with those numbers.

Full marks for a description like this need three things: a list of the population, each member numbered; random numbers generated; and how out-of-range numbers and repeats are handled, or equivalently that the process continues until nn different members are chosen.

Explaining why a method is unsatisfactory

The most common sampling question gives a method in context and asks you to explain why it is unsatisfactory, or why it does not give a random sample. The answer is almost always that some members of the population cannot be chosen, or are more likely to be chosen than others, and that this matters because those members are likely to differ in the very thing being measured.

Criticising a sampling method
  1. Identify the population the question is really about.
  2. Ask who in that population cannot be selected by this method.
  3. Ask who is more likely to be selected, or who chooses whether to take part.
  4. Say how the excluded or favoured group is likely to differ in the quantity being measured, so the results will be too high or too low.
  5. Write all of this in the context of the question.

The usual sources of bias are worth recognising on sight.

  • Convenience: choosing whoever is nearest or easiest to reach (the first people through the door, the plants at the edge of the field).
  • Time and place: surveying at one location or one time excludes everyone who is elsewhere then.
  • Self-selection: when people choose whether to respond (phone-ins, online polls, a questionnaire left on a desk), those with strong opinions are over-represented.
  • An incomplete sampling frame: a list that leaves out part of the population, such as a phone directory that excludes people with only mobile phones.
  • Hidden patterns: taking every 10th item from a production line can coincide with a machine fault that recurs every 10 items.
  • Non-response: if many people chosen do not reply, the replies may come from an unrepresentative group.
A survey at the station

A researcher wants to estimate the mean time that working adults in a town spend travelling to work. She interviews 5050 people at the town's railway station between 77 am and 88 am on a Monday.

(a) Give two reasons why this sample is unlikely to be representative.

(b) Suggest a better method.

Solution

(a) Only people who travel by train can be chosen, so people who drive, cycle, walk or work from home are excluded. Train journeys are likely to be longer than many of these, so the estimate is likely to be too high.

Only people travelling between 77 am and 88 am on a Monday can be chosen, so people with different working hours, or who do not work on Mondays, are excluded.

(b) Obtain a list of all working adults in the town (for example from a register of residents), number them, and use random numbers to choose 5050 of them to ask.

Comparing three methods

A council wants to estimate the proportion of households in a town that would use a new recycling service. Three methods are suggested.

A: Phone 200200 numbers chosen at random from the local landline directory.

B: Ask every household in one street, chosen at random.

C: Put a notice in the local newspaper asking households to reply online.

For each method, give one reason why it may not give a representative sample.

Solution

A: Households without a landline, or whose number is not listed, cannot be chosen. These may be younger or less settled households, whose views on recycling may differ.

B: Households in one street tend to be similar (similar housing, income, and the same existing collection arrangements), so their views will not represent the whole town. Every household in the town does have a chance of selection, but not every set of households: only sets that make up a single street can occur.

C: The sample is self-selected. Only households that read the paper can respond, and of those, the households most interested in recycling are the most likely to reply, so the proportion is likely to be over-estimated.

Every member equally likely is not enough

The definition of a random sample asks for more than "every member has an equal chance". Every possible sample of size nn must be equally likely. Some methods pass the first test and fail the second.

A subtle failure of randomness

A school has 600600 students listed alphabetically in a register, 2020 students to a page over 3030 pages. To choose a sample of 2020 students, a teacher picks one page at random and takes all 2020 students on that page.

(a) Show that every student has the same probability of being chosen.

(b) Explain why this is nevertheless not a random sample of size 2020.

(c) Explain why the sample might be biased in a survey about family background.

Solution

(a) A student is chosen exactly when their page is chosen, which happens with probability 130\dfrac{1}{30}. This is the same for every student, and 130=20600\dfrac{1}{30} = \dfrac{20}{600}, the same as in a genuinely random sample.

(b) Only 3030 different samples are possible: the 3030 pages. A sample containing, say, the first student on page 11 and the last student on page 3030 can never occur. In a random sample every set of 2020 students would be possible, so this method is not random.

(c) Students on one page have surnames beginning with the same few letters. Students who share a surname may be siblings, and surnames are linked to family and cultural background, so the sample could over-represent particular families or backgrounds.

A statistic is a random variable

Take two random samples of 5050 bulbs and you will get two different sample means. Before the sample is taken, the sample mean is uncertain: it is a random variable, written with a capital letter, Xˉ\bar{X}. Once a particular sample has been measured, its mean is a number, written xˉ\bar{x}.

Because Xˉ\bar{X} is a random variable, it has a distribution, with its own mean and variance. That distribution, the sampling distribution of the mean, tells you how far a sample mean is likely to be from the population mean, and is the subject of the distribution of the sample mean. The same idea gives unbiased estimates of μ\mu and σ2\sigma^2.

Common mistakes
  • Writing "it is not random" and stopping. That earns nothing. Name who is excluded or over-represented, in context, and say how this might affect the result.
  • Thinking a larger sample removes bias. A biased method gives a biased result whatever the sample size. Size reduces random variation, not bias.
  • Confusing "random" with "haphazard". Picking items "without thinking" or "the ones that look typical" is not random. Random selection uses random numbers or an equivalent physical process.
  • Forgetting repeats and out-of-range numbers. A description of random number sampling must say what to do with them.
  • Confusing a parameter with a statistic. μ\mu is a fixed (unknown) population value; xˉ\bar{x} is calculated from a sample and changes from sample to sample.
  • Assuming equal chances for individuals is enough. A random sample needs every possible sample to be equally likely.
Exam tip
  • "Explain why the method is unsatisfactory" is usually 1 or 2 marks. One clear, contextual reason per mark: "only people who travel by train can be chosen, so people who drive are excluded".
  • "Describe how to use random numbers" needs: number the population (from a list), generate random numbers, ignore numbers out of range and repeats, continue until nn members are chosen.
  • "Why is a sample used rather than a census?" Give a reason that fits the context: testing is destructive, or the population is too large.
  • "State what is meant by a population parameter" or "a random sample": use the definitions above.
  • In later parts of a question, the requirement that "the sample is random" is often the assumption you are asked to state, for example before using the Central Limit Theorem or a hypothesis test.
Summary
  • Population: everything of interest. Sample: the part actually measured. Census: measure everything.
  • Parameters (μ\mu, σ2\sigma^2, pp) describe the population and are fixed; statistics (xˉ\bar{x}, s2s^2) come from samples and vary.
  • Sample when testing is destructive, or the population is too large, slow or expensive to measure.
  • A random sample makes every possible sample of size nn equally likely; it avoids bias and justifies the later calculations.
  • Random numbers: number the sampling frame, generate numbers, ignore out-of-range values and repeats, continue until nn are chosen.
  • Criticise a method by naming who cannot be chosen or is more likely to be chosen, in context, and how that distorts the result.
  • A sample mean is a random variable Xˉ\bar{X} with its own distribution.

Practice questions

Question
  1. A manufacturer of matches wants to know the proportion of its matches that light first time. Explain why a sample must be used rather than a census.
  2. State the difference between a parameter and a statistic, using the mean mass of apples from an orchard as your example.
  3. A firm has 350350 employees. Describe how to use random numbers to choose a sample of 1212 employees.
  4. A list of 6060 items is numbered 0101 to 6060. Use the random digits 7 2 1 9 5 8 1 9 0 3 4 4 6 07\,2\,1\,9\,5\,8\,1\,9\,0\,3\,4\,4\,6\,0, read in pairs from left to right, to choose 44 items.
  5. To estimate the mean number of hours of television watched per week by children in a city, a researcher asks the children at one primary school. Give two reasons why the sample may not be representative.
  6. A radio station asks listeners to phone in to say whether they support a new road. Of 400400 callers, 320320 oppose it. Explain why it would be unwise to conclude that 80%80\% of the region's population oppose the road.
  7. A quality inspector tests every 25th bottle filled by a machine that has 55 filling nozzles used in strict rotation. Explain why this method could fail to detect a faulty nozzle, or exaggerate its effect.
  8. A college has 12001200 students. Method P: put all 12001200 names in a box and draw 6060 without replacement. Method Q: number the students 11 to 12001200, choose one of the first 2020 numbers at random, and then take every 20th student after it. (a) Show that under each method, each student has probability 120\tfrac{1}{20} of being chosen. (b) Explain why only method P gives a random sample, and give a situation in which method Q could produce a biased sample.
Answers
  1. Testing whether a match lights destroys it, so a census would leave nothing to sell. (The population is also very large.)

  2. The population mean mass μ\mu of all the apples in the orchard is a parameter: a single fixed value, unknown in practice. The mean mass xˉ\bar{x} of a sample of, say, 4040 apples is a statistic: it is calculated from the sample and would be different for a different sample of 4040.

  3. Obtain a list of all employees and number them 001001 to 350350. Generate three-digit random numbers using a calculator or computer. Ignore 000000, numbers above 350350 and repeats. Continue until 1212 different numbers are obtained and choose the corresponding employees.

  4. Pairs: 72,19,58,19,03,44,6072, 19, 58, 19, 03, 44, 60. Ignore 7272 (out of range) and the second 1919 (repeat). The items chosen are 1919, 5858, 0303 and 4444.

  5. Only children at that school can be chosen; children at other schools, which may be in different areas with different family backgrounds, are excluded. Also only primary-age children are included, so older children (who may watch different amounts) are excluded.

  6. The callers are self-selected: only people who listen to that station can call, and people with strong opinions, especially opponents, are much more likely to phone. The sample is biased, so 80%80\% is likely to over-estimate the proportion who oppose the road.

  7. Since 2525 is a multiple of 55, every 25th bottle comes from the same nozzle. If that nozzle works, faults in the other four are never seen; if that nozzle is the faulty one, every tested bottle is faulty and the fault rate is greatly exaggerated.

  8. (a) P: by symmetry each student is equally likely, and 6060 of 12001200 are chosen, so the probability is 601200=120\dfrac{60}{1200} = \dfrac{1}{20}. Q: a student is chosen exactly when the starting number is the one in their position within each block of 2020, which has probability 120\dfrac{1}{20}. (b) In P every set of 6060 students is equally likely. In Q only 2020 different samples are possible (one for each starting number); for example two consecutively numbered students can never both be chosen, so Q is not a random sample. If the list were arranged in tutor groups of 2020 with, say, the group representative listed first, then Q would either pick all the representatives or none of them.

How well do you know this?

Where this leads

Console

Search notes, courses and tools, or run an action