Cumulative Frequency Graphs
A cumulative frequency graph answers the question "how many values are less than this?" for every value at once. That makes it the tool for estimating the median, the quartiles and any percentile of a large grouped data set, and for finding how many values lie above, below or between given values. The syllabus asks you to draw and interpret these graphs, and to use them "to estimate medians, quartiles, percentiles, the proportion of a distribution above (or below) a given value, or between two values". Expect a full question: build the table, draw the graph, read off estimates, then often draw a box plot or compare with another data set.
Cumulative frequency
The cumulative frequency at a value is the number of data values less than (or equal to) . For grouped data it is found at each upper class boundary by adding up the frequencies of all classes up to and including that class.
A cumulative frequency table adds a running total to the frequency table. The key idea is where each running total belongs: when you have counted all the values in the class , you know how many are less than . So each cumulative frequency is plotted at the upper class boundary, never at the mid-point.
Drawing the graph
- Find the upper class boundary of each class (see Histograms for boundaries of rounded, discrete and age data).
- Add up the frequencies to get the cumulative frequency at each upper boundary. The last one must equal the total .
- Draw axes: the variable (with units) across, cumulative frequency up, from to .
- Plot the point (lower boundary of the first class, ): no values are below the start of the first class.
- Plot (upper boundary, cumulative frequency) for each class.
- Join the points with a smooth curve or with straight line segments. Do not extend the graph beyond the last point.
Straight segments correspond to assuming the values are spread evenly within each class, the same assumption used for histograms. A smooth curve usually gives readings very close to this. Either is accepted unless the question says otherwise. The worked answers below use straight segments (linear interpolation), so they can be checked exactly; readings from your own smooth curve may differ slightly and are accepted within a tolerance.
Reading the graph
For large grouped data sets with values, read across from these cumulative frequencies.
| Quantity | Read across at cumulative frequency |
|---|---|
| Median | |
| Lower quartile | |
| Upper quartile | |
| th percentile |
- Number below a value : read up from to the curve, then across.
- Number above : minus the number below.
- Number between and : (number below ) (number below ).
For grouped data you use , not the used for raw data: the graph is a continuous model of the data, and the half-way point of the total is what you want. Every value read from the graph is an estimate.
Here are the journey times of people.
| Time (min) | ||||||
|---|---|---|---|---|---|---|
| Frequency |
| Time less than (min) | |||||||
|---|---|---|---|---|---|---|---|
| Cumulative frequency |
Calculating instead of reading: linear interpolation
When a question gives a cumulative frequency table but asks for an estimate "by calculation", or when you want to check a graph reading, use linear interpolation: assume the values in a class are spread evenly.
To estimate the value at cumulative frequency :
- Find the class whose cumulative frequencies straddle : say it runs from boundary to boundary , with cumulative frequency before it and frequency in it.
- The fraction of the way through the class is .
- Estimate .
For the journey times, is at . That lies in , with and :
Worked examples
Use the journey-time graph above to estimate the median and the interquartile range.
Solution
. Read across at for the median, for and for .
- Median: cumulative frequency is reached exactly at . Median minutes.
- : minutes.
- : lies in , , : minutes.
IQR minutes. (Readings from a smooth curve between about and minutes would be accepted.)
For the same journey times, estimate:
(a) the number of people whose journey took more than minutes; (b) the th percentile; (c) the percentage of people whose journey took between and minutes.
Solution
(a) At : halfway through , so cumulative frequency . More than minutes: people.
(b) of is . This lies in : minutes.
(c) At : . At : . Between: people, which is , about .
The marks of students in an exam (whole numbers from to ) are grouped.
| Mark | – | – | – | – | – |
|---|---|---|---|---|---|
| Frequency |
(a) State the coordinates of the points you would plot for a cumulative frequency graph. (b) The top of students get grade A. Estimate the lowest mark for grade A. (c) of students pass. Estimate the pass mark.
Solution
(a) Marks are discrete, so the upper boundaries are , and the graph starts at the lower boundary . Points: , , , , , .
(b) The top is students, so the boundary is where cumulative frequency . This lies in to : . Students scoring or more get grade A.
(c) fail: students, so read at cumulative frequency , in to : . The pass mark is about (students scoring or more pass).
In context, round grade boundaries to whole marks and say which way you rounded.
For the journey times, the shortest journey was minutes and the longest minutes. A second group of people had median minutes, , , shortest and longest minutes.
(a) Describe the box plot for the first group. (b) Compare the two groups' journey times.
Solution
(a) From the earlier results: minimum , , median , , maximum . Box from to with the median at ; whiskers to and .
(b) The second group's journeys took longer on average (median against minutes). The second group's times were less variable (IQR against about minutes; range against minutes).
The heights, cm, of seedlings are grouped as (frequency ), (frequency ), (frequency ) and (frequency ). Using linear interpolation, the median is estimated to be . Find and .
Solution
Total: , so .
The median (th value, at cumulative frequency ) is , which lies in . Before this class the cumulative frequency is . By interpolation:
Substitute : , so , and .
Check: cumulative frequency at is ; we need more out of , a fifth of the class width, giving .
Plotting at mid-points or lower boundaries. Cumulative frequency is "how many are less than", so it belongs at the upper boundary. Plotting at mid-points shifts the whole graph left by half a class.
Forgetting the starting point. The graph must start at (lowest boundary, ). Without it the first class cannot be read.
Reading "more than" as "less than". The graph gives numbers below a value. For "more than", subtract from . For percentages, divide by and multiply by .
- Show the cumulative frequency table; the plotted points earn marks independently of the curve.
- Draw your reading lines on the graph (across, then down). Examiners look for them, and they protect method marks.
- Use , and for grouped data on a cumulative frequency graph.
- Give answers as estimates, to a sensible accuracy (usually what you can read from the graph, such as to the nearest or unit).
- If a question asks for the "least mark to get a grade" or a similar practical value, give a whole number and make sure it makes sense in context.
- Cumulative frequency at an upper class boundary total of all frequencies up to that class.
- Plot (upper boundary, cumulative frequency), starting from (lowest boundary, ). Join with a smooth curve or straight lines.
- Median at , quartiles at and , th percentile at .
- Below : read up and across. Above : minus that. Between: subtract.
- Linear interpolation: .
- All values are estimates, because individual values within classes are unknown.
Practice questions
- The masses, grams, of apples are grouped as (), (), (), (), (). Construct a cumulative frequency table, and estimate the median and interquartile range.
- For the apples in question 1, estimate the number of apples lighter than and the number heavier than .
- For the journey times in this note, estimate the time exceeded by the slowest of people.
- The numbers of text messages sent by people in a day are grouped as –, –, –, – with frequencies . Write down the coordinates of the points to plot for a cumulative frequency graph.
- Explain why the median estimated from a cumulative frequency graph is only an estimate.
- The times taken by two groups of people to complete a task are grouped (minutes):
| Time | |||||
|---|---|---|---|---|---|
| Group A | |||||
| Group B |
Estimate the median and interquartile range for each group, and compare the groups. 7. The lengths, cm, of objects are grouped as (), (), (), (). The median, estimated by linear interpolation, is . Find and . 8. A newspaper claims that "more than of the people in the journey-time survey in this note took more than minutes". Use the data to decide whether the claim is justified.
Answers
-
Cumulative frequencies at : (starting from at ). Median at : . at : . at : . IQR .
-
At : apples lighter. At : , so apples heavier.
-
The slowest means the th percentile: cumulative frequency . From the earlier example, cumulative frequency is at minutes. The slowest took more than about minutes.
-
Discrete counts, so boundaries at . Points: , , , , .
-
The data are grouped, so the individual values are unknown. The graph assumes values are spread evenly (or smoothly) through each class, which is unlikely to be exactly true.
-
Group A: cumulative frequencies . Median at : . at : . at : . IQR minutes. Group B: cumulative frequencies . Median: . : . : . IQR minutes. Group A were quicker on average (median about against minutes); Group B's times were less spread out (IQR about against minutes).
-
, so . Median at in : , so , i.e. . With : , so and .
-
At minutes: cumulative frequency . More than minutes: people, which is . The claim is not justified by the data (the estimate is under ).