Cumulative Frequency Graphs

AS · S1 · 13 min

A cumulative frequency graph answers the question "how many values are less than this?" for every value at once. That makes it the tool for estimating the median, the quartiles and any percentile of a large grouped data set, and for finding how many values lie above, below or between given values. The syllabus asks you to draw and interpret these graphs, and to use them "to estimate medians, quartiles, percentiles, the proportion of a distribution above (or below) a given value, or between two values". Expect a full question: build the table, draw the graph, read off estimates, then often draw a box plot or compare with another data set.

Cumulative frequency

Definition

The cumulative frequency at a value xx is the number of data values less than (or equal to) xx. For grouped data it is found at each upper class boundary by adding up the frequencies of all classes up to and including that class.

A cumulative frequency table adds a running total to the frequency table. The key idea is where each running total belongs: when you have counted all the values in the class 20≤t<3020 \le t < 30, you know how many are less than 3030. So each cumulative frequency is plotted at the upper class boundary, never at the mid-point.

Drawing the graph

Method
  1. Find the upper class boundary of each class (see Histograms for boundaries of rounded, discrete and age data).
  2. Add up the frequencies to get the cumulative frequency at each upper boundary. The last one must equal the total nn.
  3. Draw axes: the variable (with units) across, cumulative frequency up, from 00 to nn.
  4. Plot the point (lower boundary of the first class, 00): no values are below the start of the first class.
  5. Plot (upper boundary, cumulative frequency) for each class.
  6. Join the points with a smooth curve or with straight line segments. Do not extend the graph beyond the last point.

Straight segments correspond to assuming the values are spread evenly within each class, the same assumption used for histograms. A smooth curve usually gives readings very close to this. Either is accepted unless the question says otherwise. The worked answers below use straight segments (linear interpolation), so they can be checked exactly; readings from your own smooth curve may differ slightly and are accepted within a tolerance.

Reading the graph

For large grouped data sets with nn values, read across from these cumulative frequencies.

Key result
QuantityRead across at cumulative frequency
Median12n\tfrac{1}{2} n
Lower quartile Q1Q_114n\tfrac{1}{4} n
Upper quartile Q3Q_334n\tfrac{3}{4} n
kkth percentilek100n\tfrac{k}{100} n
  • Number below a value xx: read up from xx to the curve, then across.
  • Number above xx: nn minus the number below.
  • Number between aa and bb: (number below bb) −- (number below aa).

For grouped data you use n2\tfrac{n}{2}, not the n+12\tfrac{n+1}{2} used for raw data: the graph is a continuous model of the data, and the half-way point of the total is what you want. Every value read from the graph is an estimate.

Here are the journey times of 120120 people.

Time tt (min)0≤t<100 \le t < 1010≤t<2010 \le t < 2020≤t<3020 \le t < 3030≤t<4030 \le t < 4040≤t<6040 \le t < 6060≤t<8060 \le t < 80
Frequency66181836363232202088
Time less than (min)00101020203030404060608080
Cumulative frequency0066242460609292112112120120
0 10 20 30 40 50 60 70 80 Time (minutes) 0 20 40 60 80 100 120 Cumulative frequency Q1 median Q3
Points at upper class boundaries joined by straight segments. Dashed lines read across at 30, 60 and 90 (a quarter, half and three quarters of 120).

Calculating instead of reading: linear interpolation

When a question gives a cumulative frequency table but asks for an estimate "by calculation", or when you want to check a graph reading, use linear interpolation: assume the values in a class are spread evenly.

Method

To estimate the value at cumulative frequency cc:

  1. Find the class whose cumulative frequencies straddle cc: say it runs from boundary LL to boundary UU, with cumulative frequency FF before it and frequency ff in it.
  2. The fraction of the way through the class is c−Ff\dfrac{c - F}{f}.
  3. Estimate =L+c−Ff×(U−L)= L + \dfrac{c - F}{f} \times (U - L).

For the journey times, Q1Q_1 is at c=30c = 30. That lies in 20≤t<3020 \le t < 30, with F=24F = 24 and f=36f = 36:

Q1≈20+30−2436×10=21.7 minutes.Q_1 \approx 20 + \frac{30 - 24}{36} \times 10 = 21.7\ \text{minutes}.

Worked examples

Median and quartiles from the graph (routine)

Use the journey-time graph above to estimate the median and the interquartile range.

Solution

n=120n = 120. Read across at 6060 for the median, 3030 for Q1Q_1 and 9090 for Q3Q_3.

  • Median: cumulative frequency 6060 is reached exactly at t=30t = 30. Median ≈30\approx 30 minutes.
  • Q1Q_1: 20+30−2436×10=21.720 + \tfrac{30 - 24}{36} \times 10 = 21.7 minutes.
  • Q3Q_3: 9090 lies in 30≤t<4030 \le t < 40, F=60F = 60, f=32f = 32: 30+90−6032×10=39.430 + \tfrac{90 - 60}{32} \times 10 = 39.4 minutes.

IQR ≈39.4−21.7=17.7\approx 39.4 - 21.7 = 17.7 minutes. (Readings from a smooth curve between about 1717 and 1919 minutes would be accepted.)

Above, between and percentiles

For the same 120120 journey times, estimate:

(a) the number of people whose journey took more than 5050 minutes; (b) the 9090th percentile; (c) the percentage of people whose journey took between 1515 and 3535 minutes.

Solution

(a) At t=50t = 50: halfway through 40≤t<6040 \le t < 60, so cumulative frequency ≈92+1020×20=102\approx 92 + \tfrac{10}{20} \times 20 = 102. More than 5050 minutes: 120−102=18120 - 102 = 18 people.

(b) 90%90\% of 120120 is 108108. This lies in 40≤t<6040 \le t < 60: 40+108−9220×20=5640 + \tfrac{108 - 92}{20} \times 20 = 56 minutes.

(c) At t=15t = 15: 6+510×18=156 + \tfrac{5}{10} \times 18 = 15. At t=35t = 35: 60+510×32=7660 + \tfrac{5}{10} \times 32 = 76. Between: 76−15=6176 - 15 = 61 people, which is 61120×100=50.8%\tfrac{61}{120} \times 100 = 50.8\%, about 51%51\%.

Discrete data and grade boundaries (exam style)

The marks of 200200 students in an exam (whole numbers from 00 to 100100) are grouped.

Mark00–19192020–39394040–59596060–79798080–100100
Frequency14143636707056562424

(a) State the coordinates of the points you would plot for a cumulative frequency graph. (b) The top 20%20\% of students get grade A. Estimate the lowest mark for grade A. (c) 85%85\% of students pass. Estimate the pass mark.

Solution

(a) Marks are discrete, so the upper boundaries are 19.5,39.5,59.5,79.5,100.519.5, 39.5, 59.5, 79.5, 100.5, and the graph starts at the lower boundary −0.5-0.5. Points: (−0.5,0)(-0.5, 0), (19.5,14)(19.5, 14), (39.5,50)(39.5, 50), (59.5,120)(59.5, 120), (79.5,176)(79.5, 176), (100.5,200)(100.5, 200).

(b) The top 20%20\% is 4040 students, so the boundary is where cumulative frequency =160= 160. This lies in 59.559.5 to 79.579.5: 59.5+160−12056×20=73.859.5 + \tfrac{160 - 120}{56} \times 20 = 73.8. Students scoring 7474 or more get grade A.

(c) 15%15\% fail: 3030 students, so read at cumulative frequency 3030, in 19.519.5 to 39.539.5: 19.5+30−1436×20=28.419.5 + \tfrac{30 - 14}{36} \times 20 = 28.4. The pass mark is about 2929 (students scoring 2929 or more pass).

In context, round grade boundaries to whole marks and say which way you rounded.

From the graph to a box plot, and a comparison

For the 120120 journey times, the shortest journey was 22 minutes and the longest 7878 minutes. A second group of 120120 people had median 3434 minutes, Q1=28Q_1 = 28, Q3=41Q_3 = 41, shortest 1212 and longest 7070 minutes.

(a) Describe the box plot for the first group. (b) Compare the two groups' journey times.

Solution

(a) From the earlier results: minimum 22, Q1≈21.7Q_1 \approx 21.7, median ≈30\approx 30, Q3≈39.4Q_3 \approx 39.4, maximum 7878. Box from 21.721.7 to 39.439.4 with the median at 3030; whiskers to 22 and 7878.

(b) The second group's journeys took longer on average (median 3434 against 3030 minutes). The second group's times were less variable (IQR 41−28=1341 - 28 = 13 against about 17.717.7 minutes; range 5858 against 7676 minutes).

Unknown frequencies from an estimated median (exam-hard)

The heights, hh cm, of 8080 seedlings are grouped as 0≤h<100 \le h < 10 (frequency 1212), 10≤h<2010 \le h < 20 (frequency pp), 20≤h<3020 \le h < 30 (frequency qq) and 30≤h<5030 \le h < 50 (frequency 1616). Using linear interpolation, the median is estimated to be 22 cm22\ \text{cm}. Find pp and qq.

Solution

Total: 12+p+q+16=8012 + p + q + 16 = 80, so p+q=52p + q = 52.

The median (4040th value, at cumulative frequency 4040) is 2222, which lies in 20≤h<3020 \le h < 30. Before this class the cumulative frequency is 12+p12 + p. By interpolation:

20+40−(12+p)q×10=22⇒28−pq=0.2⇒28−p=0.2q.20 + \frac{40 - (12 + p)}{q} \times 10 = 22 \quad\Rightarrow\quad \frac{28 - p}{q} = 0.2 \quad\Rightarrow\quad 28 - p = 0.2q.

Substitute p=52−qp = 52 - q: 28−52+q=0.2q28 - 52 + q = 0.2q, so 0.8q=240.8q = 24, q=30q = 30 and p=22p = 22.

Check: cumulative frequency at 2020 is 3434; we need 66 more out of 3030, a fifth of the class width, giving 2222.

Watch out

Plotting at mid-points or lower boundaries. Cumulative frequency is "how many are less than", so it belongs at the upper boundary. Plotting at mid-points shifts the whole graph left by half a class.

Watch out

Forgetting the starting point. The graph must start at (lowest boundary, 00). Without it the first class cannot be read.

Watch out

Reading "more than" as "less than". The graph gives numbers below a value. For "more than", subtract from nn. For percentages, divide by nn and multiply by 100100.

Exam tip
  • Show the cumulative frequency table; the plotted points earn marks independently of the curve.
  • Draw your reading lines on the graph (across, then down). Examiners look for them, and they protect method marks.
  • Use n2\tfrac{n}{2}, n4\tfrac{n}{4} and 3n4\tfrac{3n}{4} for grouped data on a cumulative frequency graph.
  • Give answers as estimates, to a sensible accuracy (usually what you can read from the graph, such as to the nearest 0.50.5 or 11 unit).
  • If a question asks for the "least mark to get a grade" or a similar practical value, give a whole number and make sure it makes sense in context.
Summary
  • Cumulative frequency at an upper class boundary == total of all frequencies up to that class.
  • Plot (upper boundary, cumulative frequency), starting from (lowest boundary, 00). Join with a smooth curve or straight lines.
  • Median at n2\tfrac{n}{2}, quartiles at n4\tfrac{n}{4} and 3n4\tfrac{3n}{4}, kkth percentile at k100n\tfrac{k}{100}n.
  • Below xx: read up and across. Above xx: nn minus that. Between: subtract.
  • Linear interpolation: L+c−Ff(U−L)L + \tfrac{c - F}{f}(U - L).
  • All values are estimates, because individual values within classes are unknown.

Practice questions

Question
  1. The masses, mm grams, of 100100 apples are grouped as 100≤m<120100 \le m < 120 (1010), 120≤m<130120 \le m < 130 (2222), 130≤m<140130 \le m < 140 (3535), 140≤m<150140 \le m < 150 (2121), 150≤m<170150 \le m < 170 (1212). Construct a cumulative frequency table, and estimate the median and interquartile range.
  2. For the apples in question 1, estimate the number of apples lighter than 125 g125\ \text{g} and the number heavier than 155 g155\ \text{g}.
  3. For the journey times in this note, estimate the time exceeded by the slowest 15%15\% of people.
  4. The numbers of text messages sent by 8080 people in a day are grouped as 00–99, 1010–1919, 2020–2929, 3030–5959 with frequencies 10,24,30,1610, 24, 30, 16. Write down the coordinates of the points to plot for a cumulative frequency graph.
  5. Explain why the median estimated from a cumulative frequency graph is only an estimate.
  6. The times taken by two groups of 5050 people to complete a task are grouped (minutes):
Time0≤t<50 \le t < 55≤t<105 \le t < 1010≤t<1510 \le t < 1515≤t<2015 \le t < 2020≤t<3020 \le t < 30
Group A4410101616121288
Group B2266141418181010

Estimate the median and interquartile range for each group, and compare the groups. 7. The lengths, xx cm, of 6060 objects are grouped as 0≤x<200 \le x < 20 (88), 20≤x<3020 \le x < 30 (aa), 30≤x<4030 \le x < 40 (bb), 40≤x<6040 \le x < 60 (1010). The median, estimated by linear interpolation, is 32 cm32\ \text{cm}. Find aa and bb. 8. A newspaper claims that "more than 40%40\% of the people in the journey-time survey in this note took more than 3535 minutes". Use the data to decide whether the claim is justified.

Answers
  1. Cumulative frequencies at 120,130,140,150,170120, 130, 140, 150, 170: 10,32,67,88,10010, 32, 67, 88, 100 (starting from 00 at 100100). Median at 5050: 130+50−3235×10=135.1 g130 + \tfrac{50 - 32}{35} \times 10 = 135.1\ \text{g}. Q1Q_1 at 2525: 120+25−1022×10=126.8 g120 + \tfrac{25 - 10}{22} \times 10 = 126.8\ \text{g}. Q3Q_3 at 7575: 140+75−6721×10=143.8 g140 + \tfrac{75 - 67}{21} \times 10 = 143.8\ \text{g}. IQR ≈143.8−126.8=17.0 g\approx 143.8 - 126.8 = 17.0\ \text{g}.

  2. At 125125: 10+510×22=2110 + \tfrac{5}{10} \times 22 = 21 apples lighter. At 155155: 88+520×12=9188 + \tfrac{5}{20} \times 12 = 91, so 100−91=9100 - 91 = 9 apples heavier.

  3. The slowest 15%15\% means the 8585th percentile: cumulative frequency 0.85×120=1020.85 \times 120 = 102. From the earlier example, cumulative frequency 102102 is at 5050 minutes. The slowest 15%15\% took more than about 5050 minutes.

  4. Discrete counts, so boundaries at −0.5,9.5,19.5,29.5,59.5-0.5, 9.5, 19.5, 29.5, 59.5. Points: (−0.5,0)(-0.5, 0), (9.5,10)(9.5, 10), (19.5,34)(19.5, 34), (29.5,64)(29.5, 64), (59.5,80)(59.5, 80).

  5. The data are grouped, so the individual values are unknown. The graph assumes values are spread evenly (or smoothly) through each class, which is unlikely to be exactly true.

  6. Group A: cumulative frequencies 4,14,30,42,504, 14, 30, 42, 50. Median at 2525: 10+25−1416×5=13.410 + \tfrac{25 - 14}{16} \times 5 = 13.4. Q1Q_1 at 12.512.5: 5+12.5−410×5=9.255 + \tfrac{12.5 - 4}{10} \times 5 = 9.25. Q3Q_3 at 37.537.5: 15+37.5−3012×5=18.115 + \tfrac{37.5 - 30}{12} \times 5 = 18.1. IQR ≈8.9\approx 8.9 minutes. Group B: cumulative frequencies 2,8,22,40,502, 8, 22, 40, 50. Median: 15+25−2218×5=15.815 + \tfrac{25 - 22}{18} \times 5 = 15.8. Q1Q_1: 10+12.5−814×5=11.610 + \tfrac{12.5 - 8}{14} \times 5 = 11.6. Q3Q_3: 15+37.5−2218×5=19.315 + \tfrac{37.5 - 22}{18} \times 5 = 19.3. IQR ≈7.7\approx 7.7 minutes. Group A were quicker on average (median about 13.413.4 against 15.815.8 minutes); Group B's times were less spread out (IQR about 7.77.7 against 8.98.9 minutes).

  7. 8+a+b+10=608 + a + b + 10 = 60, so a+b=42a + b = 42. Median at 3030 in 30≤x<4030 \le x < 40: 30+30−(8+a)b×10=3230 + \tfrac{30 - (8 + a)}{b} \times 10 = 32, so 22−ab=0.2\tfrac{22 - a}{b} = 0.2, i.e. 22−a=0.2b22 - a = 0.2b. With a=42−ba = 42 - b: −20+b=0.2b-20 + b = 0.2b, so b=25b = 25 and a=17a = 17.

  8. At 3535 minutes: cumulative frequency ≈60+510×32=76\approx 60 + \tfrac{5}{10} \times 32 = 76. More than 3535 minutes: 120−76=44120 - 76 = 44 people, which is 44120=36.7%\tfrac{44}{120} = 36.7\%. The claim is not justified by the data (the estimate is under 40%40\%).

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action