Types of Data and Choosing a Diagram
Before you can draw a diagram or calculate an average, you have to know what kind of data you are holding. The type of data decides which diagrams are allowed, and the question you want answered decides which of those is best. Cambridge tests this directly: the first syllabus outcome for Paper 5 asks you to "select a suitable way of presenting raw statistical data, and discuss advantages and/or disadvantages that particular representations may have". These are usually one- or two-mark parts, and they are lost by vague answers far more often than by wrong ones.
Kinds of data
Data are the values recorded for each member of a sample or population. Each thing being measured or counted is a variable.
- Qualitative (or categorical) data are descriptions, not numbers: eye colour, type of car, favourite sport.
- Quantitative data are numerical: heights, scores, numbers of goals.
- Discrete quantitative data can take only particular separate values, usually whole numbers, found by counting: the number of children in a family, the score on a die, shoe sizes.
- Continuous quantitative data can take any value in a range, found by measuring: time, length, mass, temperature.
A useful test: if you can sensibly ask "what is between 3 and 4?" and get a meaningful answer, the data are continuous. There is no family with 3.6 children, but there is a run that takes 3.6 minutes.
Two traps catch students every year.
- Continuous data are always rounded when recorded. A mass written as to the nearest gram could be anything from up to . The data are still continuous, even though every recorded value is a whole number. This matters when you work out class boundaries for histograms and cumulative frequency graphs.
- Age is continuous, but recorded oddly. People say they are until the day they turn , so "age " means . A class "– years" therefore runs from to , not from to .
Raw and grouped data
Raw data are the individual values exactly as recorded, for example the marks scored by a class. Grouped data have been sorted into classes with a frequency for each class, for example ": students".
Grouping makes a large data set manageable, but it destroys information. Once the marks are grouped you no longer know the actual marks, so every mean, median or quartile you calculate from grouped data is only an estimate. This is one of the most frequently examined "explain" points: say estimate, and say why (the exact values within each class are not known).
The four syllabus diagrams
The Paper 5 syllabus names four diagrams. Each has its own note; here is what each one is for.
| Diagram | Best for | Keeps individual values? | Shows |
|---|---|---|---|
| Stem-and-leaf | Small raw data sets (roughly 10 to 50 values), discrete or continuous; back-to-back for comparing two sets | Yes | Shape, all values, median and quartiles can be read off |
| Box-and-whisker | Comparing two or more data sets quickly; any size | No | Median, quartiles, range, IQR, skew |
| Histogram | Large data sets of continuous (grouped) data, especially with unequal class widths | No | Shape of the distribution, modal class; area shows frequency |
| Cumulative frequency graph | Large grouped data sets where you need the median, quartiles, percentiles, or the number above or below a value | No | Estimates of median, quartiles, percentiles; proportions above, below or between values |
Bar charts and pie charts are not named in the Paper 5 syllabus, but you may still be asked to compare them with a syllabus diagram. A bar chart has gaps between the bars and suits qualitative or discrete data; a histogram has no gaps because the variable is continuous and area represents frequency.
Choosing a diagram
Ask four questions, in this order.
- Is the data qualitative or quantitative? Qualitative data cannot go in any of the four syllabus diagrams. Use a bar chart or a pie chart.
- Raw or grouped? A stem-and-leaf diagram needs the raw values. Grouped data can only go in a histogram or a cumulative frequency graph (or a box plot built from estimated quartiles).
- How many values? A stem-and-leaf diagram with leaves is unreadable. Large data sets suit histograms, cumulative frequency graphs and box plots.
- What is the diagram for? To compare two sets, use back-to-back stem-and-leaf (small, raw) or two box plots on the same scale (any size). To show the shape, use a histogram or stem-and-leaf. To estimate medians, percentiles or "how many are over ", use a cumulative frequency graph.
Advantages and disadvantages that earn marks
Examiners want a specific, relevant statement. "It is easier to read" scores nothing; "it shows the median and quartiles directly, so the two data sets can be compared at a glance" scores the mark. The statements below are the ones mark schemes accept.
Stem-and-leaf diagram
- Advantage: it shows every original data value, so nothing is lost and exact median and quartiles can be found.
- Advantage: it shows the shape of the distribution, like a histogram turned on its side.
- Disadvantage: impractical for large data sets.
- Disadvantage: harder to compare two data sets than with box plots when the sets are large.
Box-and-whisker plot
- Advantage: shows the median, quartiles and range (and so the IQR) directly.
- Advantage: very easy to compare two or more data sets when drawn on the same scale.
- Advantage: shows skew clearly.
- Disadvantage: individual values are lost; you cannot tell how many values there are or where the mode is.
- Disadvantage: hides features such as two peaks (bimodal data) or gaps.
Histogram
- Advantage: suits large sets of continuous data.
- Advantage: shows the shape of the distribution and the modal class, even with unequal class widths, because area represents frequency.
- Disadvantage: individual values are lost; the mean and median can only be estimated.
- Disadvantage: does not show the median or quartiles directly.
Cumulative frequency graph
- Advantage: gives estimates of the median, quartiles and any percentile, and the number of values above, below or between given values.
- Disadvantage: individual values are lost and the shape is harder to see.
A safe structure for any "give one advantage of A over B" question: name a feature A has that B lacks, and tie it to the context. For example: "A stem-and-leaf diagram shows the individual times, so the exact median time of the runners can be found."
Worked examples
State whether each variable is qualitative, discrete or continuous.
(a) The number of emails a person receives in a day. (b) The time taken to download a file. (c) The make of a person's phone. (d) The shoe size of a student. (e) The mass of an apple, recorded to the nearest gram.
Solution
(a) Discrete: emails are counted, so only whole numbers are possible.
(b) Continuous: time is measured and can take any value in a range.
(c) Qualitative: a make is a category, not a number.
(d) Discrete: shoe sizes take only particular values (, , , ...), with nothing in between.
(e) Continuous: mass is measured. Rounding to the nearest gram does not make it discrete; a recorded means .
For each situation, name the most suitable diagram from stem-and-leaf, box-and-whisker, histogram and cumulative frequency graph, and give a reason.
(a) The heights of trees, grouped into classes of unequal width, to show the shape of the distribution. (b) The marks of students in each of two classes, to compare the two classes while keeping every mark. (c) The journey times of commuters, grouped, to estimate the percentage who take longer than minutes.
Solution
(a) A histogram. The data set is large, continuous and grouped, and frequency density allows unequal class widths while showing the shape correctly.
(b) A back-to-back stem-and-leaf diagram. The data sets are small, and the diagram keeps every individual mark while placing the two classes side by side for comparison.
(c) A cumulative frequency graph. It allows the number of commuters taking less than minutes to be read off, and so the number (and percentage) taking longer.
The times, in seconds, taken by athletes to run are given below.
(a) Give one advantage of representing these data in a stem-and-leaf diagram rather than a box-and-whisker plot. (b) Give one advantage of a box-and-whisker plot over a stem-and-leaf diagram if these times were to be compared with those of a second group of athletes.
Solution
(a) A stem-and-leaf diagram keeps the individual times, so the exact times (for example the fastest, seconds) can still be read off, and the shape of the distribution can be seen.
(b) A box-and-whisker plot shows the median and quartiles of each group directly, so when the two plots are drawn on the same scale the average times and the spreads of the two groups can be compared at a glance.
Note that each answer refers to a specific feature (individual values; median and quartiles) and to the context (times, groups). A bare "it is clearer" would not score.
A teacher records the number of siblings of each of the students in her class. The values are and . She wants to illustrate the data.
(a) Explain why a histogram is not a suitable diagram. (b) Suggest a more suitable diagram.
Solution
(a) The number of siblings is discrete: only whole numbers are possible. A histogram treats the variable as continuous, with bars joined across class boundaries, which suggests values such as siblings exist. (Also there are only six distinct values, so grouping would lose rather than gain clarity.)
(b) A bar chart (or vertical line graph), with one separated bar for each number of siblings. If the data are to be compared with another class, a box-and-whisker plot of each class on the same scale would also be suitable.
The ages of people at a concert are summarised below.
| Age (years) | – | – | – | – |
|---|---|---|---|---|
| Frequency |
A student draws a diagram with four touching bars whose heights are the frequencies.
(a) Explain why this diagram is misleading. (b) Find the class boundaries and the heights the bars should have in a correctly drawn histogram. (c) State which age group is most densely represented, and explain why this is not the class with the largest frequency.
Solution
(a) The classes have unequal widths. In a histogram the area of each bar represents frequency, so using frequency as height makes the wide classes (– and –) look far more important than they are.
(b) Age is recorded in completed years, so "–" means . The boundaries are .
| Age (years) | Boundaries | Width | Frequency | Frequency density |
|---|---|---|---|---|
| – | to | |||
| – | to | |||
| – | to | |||
| – | to |
The bar heights are the frequency densities , , and (people per year of age).
(c) The – class has the greatest frequency density, people per year of age, so ages are most concentrated there. The – class has the largest frequency (), but it is spread over years, giving only people per year.
Saying "discrete" for rounded measurements. Masses to the nearest gram, times to the nearest second and lengths to the nearest centimetre are still continuous. Rounding changes how the data are recorded, not what kind of quantity they are.
Generic advantages. "It is easy to read", "it looks clearer" and "it is more accurate" earn nothing. Name the feature (individual values retained, median and quartiles shown, shape shown, suitable for large data sets) and link it to the data in the question.
- "State" or "give" one advantage: one sentence, one specific feature, in context. Do not give a list hoping one item is right; a wrong statement alongside a right one can cost the mark.
- "Explain why" a diagram is unsuitable: name the property of the data (discrete, qualitative, too many values, grouped so raw values unknown) and say what goes wrong.
- When a question says "estimate" for grouped data, use the word estimate in your answer too, and if asked why, say the exact values within each class are unknown.
- If a question asks you to compare two data sets, the expected diagram is almost always two box plots on the same scale or a back-to-back stem-and-leaf diagram.
- Qualitative data are categories; quantitative data are numbers, either discrete (counted, separate values) or continuous (measured, any value in a range).
- Rounded measurements are still continuous. Age in completed years: "–" means .
- Grouping loses the individual values, so any average from grouped data is an estimate.
- Stem-and-leaf: small raw data sets; keeps every value; back-to-back for comparing two sets.
- Box-and-whisker: shows median, quartiles and range; best for comparing sets; loses individual values.
- Histogram: large grouped continuous data; area represents frequency; shows shape.
- Cumulative frequency graph: estimates median, quartiles, percentiles and numbers above or below a value.
- Every advantage or disadvantage must name a specific feature and refer to the context.
Practice questions
- Classify each as qualitative, discrete or continuous: (a) the colour of a car; (b) the number of pages in a book; (c) the temperature at noon, recorded to the nearest ; (d) the number of goals scored in a match; (e) the length of a phone call, recorded to the nearest second.
- A café records the time, in minutes, that each of customers waits for their order. Suggest a suitable diagram for displaying these times and give a reason.
- The masses of newborn babies are recorded in classes of width from to . Suggest two different suitable diagrams for different purposes, and say what each would be used for.
- Give one advantage and one disadvantage of using a box-and-whisker plot to illustrate the heights of plants.
- The waiting times of patients at two clinics are recorded, with patients at each clinic. Give one reason why a back-to-back stem-and-leaf diagram would be suitable, and one feature of the data that it would allow you to see that a pair of box plots would not.
- A newspaper shows the populations of five countries in a histogram. Explain why this is not appropriate, and suggest a better diagram.
- The ages of members of a sports club are grouped as –, –, – and –. Write down the class boundaries and the class widths.
- The lengths of fish are grouped in a table. A student says "the median length of the fish is exactly ", having found it from a cumulative frequency graph. Explain what is wrong with this statement and rewrite it correctly.
Answers
-
(a) Qualitative. (b) Discrete (counted). (c) Continuous (measured; rounding does not change this). (d) Discrete. (e) Continuous.
-
A stem-and-leaf diagram: the data set is small ( values), and the diagram keeps every individual waiting time while showing the shape of the distribution. (A box-and-whisker plot is also acceptable if the reason given is to show median and quartiles.)
-
A histogram, to show the shape of the distribution of masses (the data are continuous, grouped and numerous). A cumulative frequency graph, to estimate the median, quartiles or percentiles, or the number of babies above or below a given mass (for example the number lighter than ).
-
Advantage: it shows the median, quartiles and range of the heights directly, so the average height and spread can be read off (and it could be compared easily with another group). Disadvantage: the individual heights are lost (and features such as the mode or gaps are not shown).
-
It is suitable because each data set is small ( values) and the two sets can be placed side by side for comparison. Unlike box plots, it retains every individual waiting time, so you can see, for example, the exact longest wait, any repeated values, or whether the distribution has two peaks.
-
Countries are qualitative categories, not values on a continuous scale, so there is no meaningful horizontal axis and bars should not touch. A bar chart (with gaps between the bars) is appropriate.
-
Age in completed years, so the boundaries are . Widths: .
-
The data are grouped, so the actual lengths are not known, and a cumulative frequency graph is drawn by assuming values are spread through each class; the median can only be estimated. Correct statement: "An estimate of the median length of the fish is ."