Modelling with the Poisson Distribution

A2 · S2 · 16 min

Knowing the Poisson formula is only half the topic. The syllabus also asks you to understand why the Poisson distribution fits random events, and to use it as a model: to say when it is appropriate, to criticise it when it is not, and to check it against data. These are the "explain", "state an assumption" and "comment" parts of a Paper 6 question. They are worth one or two marks each, and they are lost far more often than the calculation marks, because students answer in generalities instead of in context.

What makes a count Poisson

Recall where the distribution comes from: chop an interval into tiny pieces, each of which contains an event with a tiny, equal probability, independently of the others. For that picture to describe reality, four things must be true.

Conditions for a Poisson model

The number of events in a fixed interval of time or space has a Poisson distribution if the events occur

  • singly: two events cannot happen at exactly the same moment or place;
  • independently: one event happening does not make another more or less likely;
  • at random: the timing or position of each event cannot be predicted;
  • at a constant average rate: the mean number of events per unit interval does not change over the interval (it is uniform).

The mean λ\lambda then follows from the rate and the length of the interval. Nothing else is needed: there is no "number of trials" and no "probability of success".

How each condition fails

Thinking about how a condition can break is the fastest way to answer "why might this model be unsuitable?".

ConditionTypical failureExample
SinglyEvents come in clustersPeople arrive at a cinema in couples and groups
IndependentlyOne event triggers othersA road accident causes a tailback, causing further accidents
Constant rateThe rate varies with time or placeCustomers at a café peak at lunchtime; traffic peaks in the rush hour
At randomEvents are scheduledBuses arrive at a stop according to a timetable

A count with a fixed maximum is also a warning sign. "The number of the 30 pupils in a class who are absent" can never exceed 30, so it is binomial (or something else), not Poisson. A Poisson variable can, in principle, take any whole-number value.

Judging whether a Poisson model is suitable

For each of the following, state whether a Poisson distribution is likely to be a suitable model, giving a reason in context.

(a) The number of misprints on a page of a newspaper.

(b) The number of customers entering a supermarket in a 5-minute period, observed at different times throughout a Saturday.

(c) The number of sixes obtained when a fair die is thrown 20 times.

(d) The number of emissions per second from a radioactive source.

(e) The number of passengers getting off a train at a station each time it stops.

Solution

(a) Suitable. Misprints occur singly and at random through the text, and one misprint does not cause another, at a roughly constant average rate per page.

(b) Not suitable as it stands. The average rate of arrivals changes during the day (busier at some times than others), so the rate is not constant. Customers may also arrive together in families, so not singly or independently.

(c) Not suitable. There is a fixed number of trials (2020) and a fixed probability 16\tfrac{1}{6}, so the count is B(20,16)B(20, \tfrac{1}{6}). The count cannot exceed 2020.

(d) Suitable. Radioactive decays occur singly, independently and at random, at a constant average rate (over a short period).

(e) Not suitable. Passengers often travel in groups, so they do not alight independently; and the number depends strongly on the time of day and the station.

Tip

"Constant average rate" does not mean events arrive at regular intervals. Poisson events arrive irregularly; it is only the long-run average that is constant. Regular arrivals (a bus every 10 minutes) are the opposite of Poisson.

Stating assumptions in context

When a question says "state, in context, an assumption needed for the Poisson model", the mark is for applying a condition to the situation in the question. Compare:

  • "Events occur independently." This is a generic statement and usually scores nothing.
  • "Breakdowns of the machine occur independently of each other." This scores.
  • "Calls to the help desk occur at a constant average rate throughout the day." This scores.
Answering an assumption question
  1. Pick one condition (independence and constant rate are the most natural to explain).
  2. Name the actual events from the question: flaws, calls, goals, bacteria.
  3. Name the actual interval: per metre, per hour, per millilitre.
  4. If asked why the model might fail, give a realistic mechanism in the context, not just "the condition might not hold".

Checking a model against data

The Poisson distribution has a fingerprint: its mean equals its variance. So the first check on any set of count data is to calculate the sample mean and variance and see whether they are close.

  • If xˉ≈\bar{x} \approx variance, a Poisson model is plausible.
  • If the variance is clearly larger than the mean, the data are overdispersed, which usually means events are clustered or the rate varies.
  • If the variance is clearly smaller than the mean, the events are more regular than random, which also rules out a Poisson model.

You may use either the ordinary variance ∑x2n−xˉ2\dfrac{\sum x^2}{n} - \bar{x}^2 or the unbiased estimate s2s^2 (see unbiased estimates). For the comparison it makes little difference with a reasonable sample size, but state which you have calculated.

Key result

For data from a frequency table with values xx and frequencies ff:

xˉ=∑fx∑f,variance=∑fx2∑f−xˉ2.\bar{x} = \frac{\sum fx}{\sum f}, \qquad \text{variance} = \frac{\sum fx^2}{\sum f} - \bar{x}^2.

A Poisson model is supported when the mean and variance are approximately equal; the sample mean is then used as the estimate of λ\lambda.

Testing the mean-variance fingerprint

The number of calls received by a switchboard in each of 200200 one-minute periods was recorded.

Number of calls00112233445566
Frequency333358585252333316166622

(a) Calculate the mean and variance of the number of calls per minute.

(b) Explain how your answers support the use of a Poisson model.

(c) Use a Poisson model to estimate the probability of at least 33 calls in a minute, and the probability of at most 22 calls in a two-minute period.

Solution

(a)

∑fx=0+58+104+99+64+30+12=367,xˉ=367200=1.835\sum fx = 0 + 58 + 104 + 99 + 64 + 30 + 12 = 367, \qquad \bar{x} = \frac{367}{200} = 1.835∑fx2=0+58+208+297+256+150+72=1041,variance=1041200−1.8352=1.838 (4 s.f.)\sum fx^2 = 0 + 58 + 208 + 297 + 256 + 150 + 72 = 1041, \qquad \text{variance} = \frac{1041}{200} - 1.835^2 = 1.838 \text{ (4 s.f.)}

(b) The mean (1.8351.835) and the variance (1.8381.838) are almost equal, which is a property of the Poisson distribution.

(c) Take λ=1.835\lambda = 1.835 for one minute. Let X∼Po(1.835)X \sim \text{Po}(1.835).

P(X≥3)=1−e−1.835(1+1.835+1.83522)=1−0.7212=0.279P(X \ge 3) = 1 - e^{-1.835}\left(1 + 1.835 + \frac{1.835^2}{2}\right) = 1 - 0.7212 = 0.279

For two minutes, Y∼Po(3.67)Y \sim \text{Po}(3.67):

P(Y≤2)=e−3.67(1+3.67+3.6722)=0.291P(Y \le 2) = e^{-3.67}\left(1 + 3.67 + \frac{3.67^2}{2}\right) = 0.291

Expected frequencies

A second check is to compare observed frequencies with the frequencies a Poisson model predicts. If there are NN observations, the expected frequency of the value rr is N×P(X=r)N \times P(X = r). For the switchboard data, with λ=1.835\lambda = 1.835:

Calls00112233445566
Observed333358585252333316166622
Expected31.931.958.658.653.753.732.932.915.115.15.55.51.71.7

The agreement is very close, which strongly supports the model. (A formal test of fit is not on the 9709 syllabus; a comment on how close the values are is all that is expected.)

When the data say no

The number of customers arriving at a café in each of 100100 one-minute periods gave the following results.

Number of customers0011223344
Frequency5252202014148866

(a) Find the mean and variance of the data.

(b) Comment on whether a Poisson distribution is a suitable model, suggesting a reason in context.

Solution

(a)

∑fx=20+28+24+24=96,xˉ=0.96\sum fx = 20 + 28 + 24 + 24 = 96, \qquad \bar{x} = 0.96∑fx2=20+56+72+96=244,variance=244100−0.962=1.518 (4 s.f.)\sum fx^2 = 20 + 56 + 72 + 96 = 244, \qquad \text{variance} = \frac{244}{100} - 0.96^2 = 1.518 \text{ (4 s.f.)}

(b) The variance (1.521.52) is considerably larger than the mean (0.960.96), so a Poisson model is not suitable. A likely reason is that customers arrive in groups (friends or families coming in together), so arrivals do not occur singly or independently. There are also more minutes with no customers (5252) than a Poisson model with mean 0.960.96 predicts (about 3838), which is consistent with clustering.

Using summary statistics

Exam questions often give ∑x\sum x and ∑x2\sum x^2 instead of a table. The method is the same.

Summary statistics and a probability

The numbers of chocolate chips, xx, in a random sample of 5050 biscuits are summarised by

∑x=420,∑x2=3940.\sum x = 420, \qquad \sum x^2 = 3940.

(a) Calculate unbiased estimates of the population mean and variance.

(b) Explain why these results suggest that the number of chips per biscuit may be modelled by a Poisson distribution.

(c) State, in context, an assumption required for the Poisson model to be valid.

(d) Using the model, find the probability that a randomly chosen biscuit contains fewer than 55 chips.

Solution

(a)

xˉ=42050=8.4,s2=149(3940−420250)=41249=8.41 (3 s.f.)\bar{x} = \frac{420}{50} = 8.4, \qquad s^2 = \frac{1}{49}\left(3940 - \frac{420^2}{50}\right) = \frac{412}{49} = 8.41 \text{ (3 s.f.)}

(b) The estimates of the mean (8.48.4) and variance (8.418.41) are approximately equal, as they are for a Poisson distribution.

(c) The chocolate chips are distributed independently and at random throughout the biscuit mixture (so that the average number of chips per biscuit is constant).

(d) X∼Po(8.4)X \sim \text{Po}(8.4).

P(X<5)=P(X≤4)=e−8.4(1+8.4+8.422+8.436+8.4424)P(X < 5) = P(X \le 4) = e^{-8.4}\left(1 + 8.4 + \frac{8.4^2}{2} + \frac{8.4^3}{6} + \frac{8.4^4}{24}\right)=e−8.4(1+8.4+35.28+98.784+207.4464)=0.0789= e^{-8.4}(1 + 8.4 + 35.28 + 98.784 + 207.4464) = 0.0789
Common mistakes
  • Context-free assumptions. "Events are independent" with no reference to the situation scores zero when the question says "in context".
  • Confusing "constant rate" with "regular". A Poisson process is irregular; a timetable is not Poisson.
  • Using the standard deviation instead of the variance. Compare the mean with the variance. A mean of 44 and a standard deviation of 22 is perfectly consistent with Poisson, since 22=42^2 = 4.
  • Using the Poisson for a bounded count. If there is a fixed number of trials, it is binomial.
  • Saying the data "prove" a Poisson distribution. Equal mean and variance supports the model; it cannot prove it.
Exam tip
  • "Comment on" or "explain" after calculating a mean and variance: quote both numbers and say whether they are approximately equal. One sentence earns the mark.
  • "Give a reason why the Poisson distribution may not be a suitable model": give a specific mechanism. "Customers may arrive in groups" or "the rate of calls is likely to be higher in the evening" scores; "the events may not be random" does not.
  • When a question tells you to "assume a Poisson distribution", do not argue with it in later parts. The comment is wanted only where asked.
  • Expect the model questions to be followed straight away by a calculation using the sample mean as λ\lambda, often in a rescaled interval.
Summary
  • Poisson counts arise when events occur singly, independently, at random and at a constant average rate.
  • Each condition has a typical failure: clusters, knock-on effects, peaks in the rate, or scheduling.
  • A count with a fixed maximum is not Poisson.
  • Mean ≈\approx variance in data supports a Poisson model; variance much larger than the mean suggests clustering or a varying rate.
  • Expected frequency of rr is N×P(X=r)N \times P(X = r); close agreement with observed frequencies supports the model.
  • Every assumption or criticism must name the events and the interval from the question.

Practice questions

Question
  1. State, with a reason, whether a Poisson distribution is likely to be suitable for each of the following. (a) The number of telephone calls received by a fire station in a day. (b) The number of red cards shown in a 90-minute football match. (c) The number of pupils in a class of 30 who are left-handed. (d) The number of buses arriving at a bus stop between 08

    and 08
    .

  2. The number of goals scored by a hockey team in a match has mean 2.32.3 and standard deviation 1.51.5. Comment on whether a Poisson distribution might be a suitable model.

  3. A shop records the number of customers arriving in 5-minute intervals from 09

    to 18
    on a weekday. Give two reasons in context why a Poisson distribution may not be a suitable model.

  4. The number of weeds per square metre in a field was counted in 8080 randomly chosen square metres.

    Weeds0011223344≥5\ge 5
    Frequency12122121222214148833

    Assuming the "≥5\ge 5" class consists of three squares each containing exactly 55 weeds, calculate the mean and variance, and comment.

  5. Using your mean from question 4, calculate the expected frequency of squares with exactly 22 weeds.

  6. 3030 accidents occurred on a stretch of road in 5050 weeks. Assuming a Poisson model, estimate the probability that there are at least 22 accidents in a 4-week period, and state one assumption you have made in context.

  7. The numbers of cracks, xx, in 8080 randomly chosen 10-metre lengths of pipe are summarised by ∑x=168\sum x = 168 and ∑x2=532\sum x^2 = 532. (a) Find unbiased estimates of the mean and variance and comment on the suitability of a Poisson model. (b) Using the model, find the probability that a 25-metre length contains no cracks. (c) Five separate 10-metre lengths are chosen. Find the probability that at most one of them contains no cracks.

  8. A biologist models the number of bacteria in 1 ml1\ \text{ml} of a well-stirred liquid as Po(λ)\text{Po}(\lambda). (a) Explain why the stirring is relevant to the validity of the model. (b) The probability that a 1 ml1\ \text{ml} sample contains no bacteria is 0.20.2. Find the probability that a 3 ml3\ \text{ml} sample contains at least 33 bacteria.

Answers
  1. (a) Suitable: calls arrive singly, independently and at random. (Over a whole day the rate may not be constant, which is a possible criticism.) (b) Suitable: red cards occur singly and randomly at a roughly constant average rate during a match (though one red card may change how a game is played, so independence is questionable). (c) Not suitable: a fixed number of trials (3030) and the count cannot exceed 3030; this is binomial. (d) Not suitable: buses run to a timetable, so arrivals are not random.

  2. Variance =1.52=2.25= 1.5^2 = 2.25, which is close to the mean 2.32.3, so a Poisson model may be suitable.

  3. Any two of: the arrival rate is likely to vary during the day (busier at lunchtime and after work), so it is not constant; customers may arrive in groups (families or friends), so they do not arrive singly or independently; an event such as a sale announcement may bring a rush of customers.

  4. ∑fx=21+44+42+32+15=154\sum fx = 21 + 44 + 42 + 32 + 15 = 154, so xˉ=154/80=1.925\bar{x} = 154/80 = 1.925. ∑fx2=21+88+126+128+75=438\sum fx^2 = 21 + 88 + 126 + 128 + 75 = 438, so variance =438/80−1.9252=5.475−3.7056=1.769= 438/80 - 1.925^2 = 5.475 - 3.7056 = 1.769. The mean (1.931.93) and variance (1.771.77) are fairly close, so a Poisson model is reasonable.

  5. 80×e−1.9251.92522=80×0.2704=21.680 \times e^{-1.925}\dfrac{1.925^2}{2} = 80 \times 0.2704 = 21.6 (compared with 2222 observed).

  6. Mean per week =30/50=0.6= 30/50 = 0.6, so per 4 weeks λ=2.4\lambda = 2.4, X∼Po(2.4)X \sim \text{Po}(2.4). P(X≥2)=1−e−2.4(1+2.4)=1−0.3084=0.692P(X \ge 2) = 1 - e^{-2.4}(1 + 2.4) = 1 - 0.3084 = 0.692. Assumption: accidents on this road occur independently of each other at a constant average rate (for example, the rate does not change with the seasons).

  7. (a) xˉ=168/80=2.1\bar{x} = 168/80 = 2.1; s2=179(532−168280)=179.279=2.27s^2 = \dfrac{1}{79}\left(532 - \dfrac{168^2}{80}\right) = \dfrac{179.2}{79} = 2.27. Mean and variance are close, so a Poisson model is plausible. (b) For 25 m, λ=2.5×2.1=5.25\lambda = 2.5 \times 2.1 = 5.25. P(X=0)=e−5.25=0.00525P(X = 0) = e^{-5.25} = 0.00525. (c) p=P(no cracks in 10 m)=e−2.1=0.12246p = P(\text{no cracks in 10 m}) = e^{-2.1} = 0.12246. Let N∼B(5,0.12246)N \sim B(5, 0.12246). P(N≤1)=0.877545+5(0.12246)(0.87754)4=0.5204+0.3631=0.884P(N \le 1) = 0.87754^5 + 5(0.12246)(0.87754)^4 = 0.5204 + 0.3631 = 0.884.

  8. (a) Stirring spreads the bacteria at random and uniformly through the liquid, so that the average number per ml is constant and bacteria are not clustered together. Without stirring, bacteria could clump or settle, breaking the constant-rate and independence conditions. (b) e−λ=0.2⇒λ=ln⁡5=1.6094e^{-\lambda} = 0.2 \Rightarrow \lambda = \ln 5 = 1.6094. For 3 ml, μ=3ln⁡5=4.8283\mu = 3\ln 5 = 4.8283. P(Y≥3)=1−e−4.8283(1+4.8283+4.828322)=1−0.008(17.4846)=1−0.1399=0.860P(Y \ge 3) = 1 - e^{-4.8283}\left(1 + 4.8283 + \dfrac{4.8283^2}{2}\right) = 1 - 0.008(17.4846) = 1 - 0.1399 = 0.860. (Note e−3ln⁡5=5−3=0.008e^{-3\ln 5} = 5^{-3} = 0.008 exactly.)

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action