Mode and Modal Class

AS · S1 · 9 min

The mode is the most common value. It is the simplest measure of central tendency, the only one that works for qualitative data, and the one that answers questions like "which size should a shop stock most of?". For grouped data the equivalent is the modal class, and there is a trap here that Paper 5 tests regularly: with unequal class widths, the modal class is the one with the greatest frequency density, not the greatest frequency.

The mode

Definition

The mode of a data set is the value that occurs most often. A data set with two values tied for most frequent is bimodal; if every value occurs equally often there is no useful mode.

  • 3,5,5,6,8,8,8,93, 5, 5, 6, 8, 8, 8, 9: the mode is 88.
  • 2,4,4,7,9,92, 4, 4, 7, 9, 9: bimodal, with modes 44 and 99.
  • Favourite colours: red 1212, blue 1919, green 77: the mode is blue. The mean and median make no sense here.

In a frequency table, the mode is the value with the highest frequency. Be careful to give the value, not the frequency: if 1414 families have 22 children and that is the largest frequency, the mode is 22 children, not 1414.

The modal class

When data are grouped, individual values are unknown, so there is no mode, only a modal class: the class in which values are most concentrated.

Key result
  • Equal class widths: the modal class has the highest frequency.
  • Unequal class widths: the modal class has the highest frequency density =frequencyclass width= \dfrac{\text{frequency}}{\text{class width}}, i.e. the tallest bar of the histogram.

The reason is the same as for histograms: a wide class can hold many values simply because it covers a long interval. The modal class is meant to show where the data are densest, which is measured per unit of the variable.

Properties of the mode

  • It is the only average for qualitative (categorical) data.
  • For discrete data it is always an actual data value, which is useful when only actual values make sense (shoe sizes, numbers of people).
  • It is not affected by extreme values.
  • It may not exist, or there may be more than one.
  • It ignores most of the data, and for small data sets it can be unstable: changing one value can move the mode a long way.
  • It is not used in further calculations (unlike the mean, which feeds the standard deviation).

Choosing an average

The three averages answer slightly different questions.

Key result
AverageUse whenStrengthWeakness
Meandata are roughly symmetrical, with no extreme valuesuses every value; used with the standard deviation; combines through totalsdistorted by extreme values and skew
Mediandata are skewed or have outliersnot affected by extreme valuesignores the sizes of most values
Modedata are qualitative, or the most common actual value is wantedalways a real value (discrete data); not affected by outliersmay not exist or be unique; ignores most data

When the data are skewed, the three averages separate in a predictable way:

  • Positive skew (long right tail): usually mode << median << mean.
  • Negative skew (long left tail): usually mean << median << mode.
  • Symmetrical: mean ≈\approx median ≈\approx mode.

The mean is pulled furthest towards the tail because it is the only one that responds to the actual size of the extreme values. More on interpreting and comparing averages is in Interpreting and comparing distributions.

Worked examples

Mode from raw data and a table (routine)

(a) Find the mode of 4,7,2,7,9,4,7,3,4,74, 7, 2, 7, 9, 4, 7, 3, 4, 7.

(b) The table shows the numbers of goals scored in 4040 matches. Find the mode, the median and the mean.

Goals001122334455
Matches8812121010663311
Solution

(a) 77 occurs four times, 44 three times, and every other value fewer. The mode is 77.

(b) Mode: the highest frequency is 1212, for 11 goal. The mode is 11 goal.

Median: cumulative frequencies 8,20,30,36,39,408, 20, 30, 36, 39, 40. With n=40n = 40, the median is the mean of the 2020th and 2121st values. The 2020th is 11 and the 2121st is 22, so the median is 1.51.5 goals.

Mean: ∑fx=0+12+20+18+12+5=67\sum fx = 0 + 12 + 20 + 18 + 12 + 5 = 67, so the mean is 6740=1.675\tfrac{67}{40} = 1.675 goals.

Mode << median << mean, consistent with the positive skew of the data (a tail of high-scoring matches).

Modal class with unequal widths

The lengths of 8080 fish (to the nearest cm) are grouped.

Length (cm)1010–14141515–19192020–29293030–4949
Frequency1515252530301010

State the modal class, with a reason.

Solution

Widths: 5,5,10,205, 5, 10, 20. Frequency densities: 3,5,3,0.53, 5, 3, 0.5.

The modal class is 1515–19 cm19\ \text{cm}, because it has the highest frequency density (55 fish per cm). The class 2020–2929 has more fish (3030), but they are spread over twice the width, so the fish are less concentrated there.

Which average? (exam style)

State, with a reason, which average is most appropriate in each case.

(a) The incomes of all the employees in a company, including the director. (b) The favourite flavours of ice cream of 200200 customers. (c) The masses of 5050 bags of sugar filled by a machine, which are roughly symmetrical.

Solution

(a) The median. Incomes are usually positively skewed, and the director's very high income would pull the mean above what most employees earn.

(b) The mode. The data are qualitative, so the mean and median cannot be calculated.

(c) The mean. The data are symmetrical with no reason to expect extreme values, and the mean uses every value (it can also be used with the standard deviation).

Unknown frequency with mode and mean (exam-hard)

The scores of some students are shown.

Score1122334455
Frequency33xx886622

(a) Given that the mode is 22, write down an inequality for xx. (b) Given also that the mean is 2.5752.575, find xx, and find the median.

Solution

(a) The mode is 22 only if its frequency beats every other frequency, the largest of which is 88: x>8x > 8, i.e. x≥9x \ge 9.

(b) ∑f=19+x\sum f = 19 + x and ∑fx=3+2x+24+24+10=61+2x\sum fx = 3 + 2x + 24 + 24 + 10 = 61 + 2x.

61+2x19+x=2.575⇒61+2x=48.925+2.575x⇒12.075=0.575x⇒x=21.\frac{61 + 2x}{19 + x} = 2.575 \quad\Rightarrow\quad 61 + 2x = 48.925 + 2.575x \quad\Rightarrow\quad 12.075 = 0.575x \quad\Rightarrow\quad x = 21.

This satisfies x≥9x \ge 9. Now n=40n = 40; cumulative frequencies are 3,24,32,38,403, 24, 32, 38, 40. The 2020th and 2121st values are both 22, so the median is 22.

Watch out

Giving the frequency instead of the value. "The mode is 1212" when 1212 is the number of matches with 11 goal. The mode is a value of the variable: 11 goal.

Watch out

Modal class by frequency with unequal widths. Always compute frequency densities before naming a modal class, and quote them as your reason.

Exam tip
  • "State the modal class" usually needs a reason when widths differ: "highest frequency density".
  • When asked which average to use, the answer must be justified by a property of these data: skewed, has an extreme value, is qualitative, or is symmetrical.
  • "Explain why the mode is not a suitable average here" usually means the data have no repeats, or the mode is at an extreme end, or there are two modes.
Summary
  • Mode: the most frequent value; can be bimodal or not exist.
  • In a frequency table the mode is the value with the largest frequency, not the frequency itself.
  • Modal class: highest frequency for equal widths; highest frequency density for unequal widths.
  • Mode is the only average for qualitative data and is unaffected by extreme values.
  • Positive skew: mode << median << mean. Negative skew: the reverse.
  • Choose the average that suits the data, and justify the choice from the data's shape or type.

Practice questions

Question
  1. Find the mode of 12,15,12,18,15,12,20,1512, 15, 12, 18, 15, 12, 20, 15.
  2. The shoe sizes of 4040 people are size 44 (22), 55 (55), 66 (99), 77 (1111), 88 (77), 99 (44), 1010 (22). Find the mode.
  3. The times, tt minutes, taken by 6060 people are grouped as 0≤t<50 \le t < 5 (99), 5≤t<105 \le t < 10 (1515), 10≤t<2010 \le t < 20 (2424), 20≤t<4020 \le t < 40 (1212). Find the modal class.
  4. Give one advantage and one disadvantage of the mode as a measure of central tendency.
  5. For a data set the mean is 4848, the median is 4242 and the mode is 3939. Describe the skewness of the data, with a reason.
  6. Suggest which average a car manufacturer should use to decide the most popular colour of car, and explain why.
  7. In a frequency table the values 1,2,3,41, 2, 3, 4 have frequencies 5,9,y,45, 9, y, 4. The mean is 2.52.5. Find yy, and show that the mode is 33.
  8. The ages of 5050 members of a club are grouped as 1010–1919 (88), 2020–2929 (1414), 3030–3939 (66), 4040–5959 (1212), 6060–8989 (1010). (a) Find the modal class. (b) A student says "the 4040–5959 class has the second-largest frequency, so it is the second most concentrated age group". Explain whether the student is right.
Answers
  1. 1212 and 1515 both occur three times, so the data are bimodal with modes 1212 and 1515.

  2. Size 77 (frequency 1111).

  3. Widths 5,5,10,205, 5, 10, 20; densities 1.8,3,2.4,0.61.8, 3, 2.4, 0.6. The modal class is 5≤t<105 \le t < 10 (highest frequency density, 33).

  4. Advantage (any one): unaffected by extreme values; can be used for qualitative data; is an actual data value for discrete data. Disadvantage (any one): may not exist or may not be unique; ignores most of the data; not used in further calculations.

  5. Mode << median << mean, so the data are positively skewed: a tail of large values pulls the mean above the median.

  6. The mode: colour is qualitative, so the mean and median cannot be found, and the manufacturer wants the most common colour.

  7. ∑f=18+y\sum f = 18 + y; ∑fx=5+18+3y+16=39+3y\sum fx = 5 + 18 + 3y + 16 = 39 + 3y. 39+3y18+y=2.5\tfrac{39 + 3y}{18 + y} = 2.5 gives 39+3y=45+2.5y39 + 3y = 45 + 2.5y, so 0.5y=60.5y = 6 and y=12y = 12. The frequency of 33 is then 1212, larger than 55, 99 and 44, so the mode is 33.

  8. (a) Ages: boundaries 10,20,30,40,60,9010, 20, 30, 40, 60, 90; widths 10,10,10,20,3010, 10, 10, 20, 30. Densities 0.8,1.4,0.6,0.6,0.3330.8, 1.4, 0.6, 0.6, 0.333. The modal class is 2020–2929. (b) The student is wrong. The 4040–5959 class covers 2020 years, so its density (0.60.6 members per year) is no higher than 3030–3939 and lower than 1010–1919 (0.80.8). The second most concentrated group is 1010–1919.

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action