In this topic, a value in a data set is classed as an outlier if it lies more than 1.5 x IQR below the lower quartile, or more than 1.5 x IQR above the upper quartile, where IQR is the interquartile range.
A small data set has lower quartile = 10 and upper quartile = 22.
(a)Calculate the interquartile range.(1)
(b)Use the rule given above to calculate the lower and upper boundaries for outliers.(2)
(Total for Question 1 is 3 marks)
2
A student records a five-number summary for a data set:
Minimum = 8, Lower quartile = 22, Median = 19, Upper quartile = 27, Maximum = 33
Explain, giving a reason, why this five-number summary cannot be correct.
(Total for Question 2 is 2 marks)
3
80 Year 11 students were surveyed about how many minutes, m, they spent on their phone after school on a given day. The grouped frequency table shows the results.
m (minutes)
Frequency
0 < m ≤ 20
8
20 < m ≤ 40
16
40 < m ≤ 60
24
60 < m ≤ 80
20
80 < m ≤ 100
8
100 < m ≤ 150
4
(a)Complete the cumulative frequency table below. Some values have already been filled in for you.
m (minutes)
Cumulative frequency
m ≤ 20
8
m ≤ 40
...
m ≤ 60
48
m ≤ 80
...
m ≤ 100
76
m ≤ 150
80
(2)
(b)Write down the coordinates you would plot on a cumulative frequency graph to represent the class 40 < m ≤ 60.(1)
(Total for Question 3 is 3 marks)
4
The cumulative frequency graph for the 80 students' screen-time data (Question 3) is shown below.
(a)Use the graph to estimate the median screen time.(2)
(b)Use the graph to estimate the lower quartile and the upper quartile of the screen times.(3)
(Total for Question 4 is 5 marks)
5
Using your quartile estimates from Question 4, answer the following.
(a)Calculate the interquartile range of the screen times.(1)
(b)A value is classed as an outlier if it lies more than 1.5 x IQR below the lower quartile, or more than 1.5 x IQR above the upper quartile. Calculate the lower and upper boundaries for outliers for this data.(3)
(Total for Question 5 is 4 marks)
6
The shortest screen time recorded in the sample was 3 minutes. The longest screen time recorded was 150 minutes.
(a)Explain why there cannot be any low outliers in this data set.(2)
(b)Determine, showing your reasoning, whether the longest recorded screen time of 150 minutes is an outlier.(2)
(Total for Question 6 is 4 marks)
7
The cumulative frequency graph for the screen-time data is shown again below. Use the graph, together with the upper outlier boundary of 127.5 minutes found in Question 5, to estimate the number of students in the sample whose screen time was greater than this boundary.
(Total for Question 7 is 3 marks)
8
It is also known that, excluding the outlier at 150 minutes, the next-longest screen time recorded was 118 minutes. Using the grid provided, and your answers to Questions 4, 5 and 6, draw a complete box plot for the screen-time data. Mark the outlier separately from the whisker.
(Total for Question 8 is 4 marks)
9
A PE teacher at a large secondary school wants to find out how many hours per week Year 10 students spend on organised sport, and whether this differs between boys and girls. There are 180 boys and 220 girls in Year 10. The teacher plans to select a stratified sample of 40 students.
(a)Work out how many boys and how many girls should be included in the sample.(2)
(b)State one advantage of using a stratified sample here rather than a simple random sample.(1)
(c)The teacher collects the data using a self-completed questionnaire, then groups the responses into a frequency table before drawing a cumulative frequency graph. Explain why grouping this continuous data is a necessary part of processing it before a cumulative frequency graph can be drawn.(2)
(d)Suggest one way the teacher could check whether the sample of 40 students is representative of the whole of Year 10.(1)
(Total for Question 9 is 6 marks)
10
A separate survey of 75 Year 8 students recorded their daily after-school screen time. The five-number summary for this sample is:
(a)Calculate the interquartile range for the Year 8 data.(1)
(b)Determine, using the 1.5 x IQR rule, whether the maximum value of 90 minutes is an outlier for the Year 8 data.(2)
(Total for Question 10 is 3 marks)
11
The box plots below show the Year 11 screen-time data (Question 8) and the Year 8 screen-time data (Question 10).
(a)Compare the median screen times of the two year groups.(1)
(b)Compare the spread of the two data sets using the interquartile range.(1)
(c)The Year 11 data set contains an outlier at 150 minutes, but the Year 8 data set does not contain any outliers. Comment on what this suggests about the two year groups.(2)
(Total for Question 11 is 4 marks)
12
Using the box plot for the Year 11 screen-time data shown below, determine whether the distribution of screen times is skewed, and justify your answer. Comment on what effect the outlier at 150 minutes has on the mean compared with the median.
(Total for Question 12 is 3 marks)
13
For a set of delivery times, the lower quartile is 20 minutes. The upper boundary for outliers, found using the 1.5 x IQR rule, is 50 minutes.
(a)Work out the upper quartile.(3)
(b)Hence state the interquartile range.(1)
(Total for Question 13 is 4 marks)
14
A quality-control inspector weighs a sample of 50 bags of crisps from a factory. For this sample, the lower quartile of the masses is 24 g and the upper quartile is 31 g.
(a)Calculate the interquartile range of the masses.(1)
(b)Use the 1.5 x IQR rule to calculate the lower and upper boundaries for outliers.(2)
(c)Five bags were set aside for closer inspection, with masses 12 g, 22 g, 28 g, 33 g and 44 g. Determine, showing your reasoning, which of these bags (if any) contain an outlier mass.(2)
(Total for Question 14 is 5 marks)
15
A shop manager records the amount, in GBP, spent by 40 customers in one hour. Most customers spent between 5 and 25 pounds, but two customers made much larger purchases of 180 pounds and 210 pounds, which are clear outliers. The manager wants to summarise the 'typical' spend and the spread of spending.
(a)State, giving a reason, whether the mean and standard deviation, or the median and interquartile range, would be more appropriate summary statistics for this data set.(2)
(b)Explain what would happen to the mean if the two outlying purchases were removed from the data set, compared with what would happen to the median.(2)
(Total for Question 15 is 4 marks)
16
A clinic records the resting heart rate, r beats per minute (bpm), of 120 patients. The grouped frequency table and the cumulative frequency graph for this data are shown below.
r (bpm)
Frequency
40 < r ≤ 50
6
50 < r ≤ 60
22
60 < r ≤ 70
38
70 < r ≤ 80
32
80 < r ≤ 90
14
90 < r ≤ 120
8
(a)Use the cumulative frequency graph to find estimates for the median and the interquartile range of the resting heart rates.(4)
(b)Calculate the lower and upper boundaries for outliers using the 1.5 x IQR rule.(2)
(Total for Question 16 is 6 marks)
17
Two additional patients, seen in the emergency department rather than as part of the routine 120-patient clinic sample above, had resting heart rates of 30 bpm and 118 bpm.
(a)Using your boundaries from Question 16(b), determine whether each of these two readings would be classed as an outlier.(2)
(b)Within the routine 120-patient sample itself (not including the two emergency department patients), the lowest heart rate recorded was 44 bpm and the highest was 96 bpm. Using the grid provided, draw a box plot for the 120-patient sample.(3)
(c)Evaluate whether it would be appropriate to include the two emergency department readings in the same data set as the 120 routine clinic patients when analysing 'normal' resting heart rates, referring to the statistical enquiry cycle.(2)
(Total for Question 17 is 7 marks)
Mark scheme · H21 Box Plots from Cumulative Frequency and Outliers
Question 1
(a) B1 12 cao
(a) Answer: 12
(b) M1 1.5 x 12 = 18, oe
(b) A1 lower boundary = -8 and upper boundary = 40 cao
B1 identifies that the median (19) is smaller than the lower quartile (22), which is impossible
B1 correct reasoning, e.g. by definition the lower quartile can never be greater than the median, since at least 25% of the data lies at or below the lower quartile and the median is the middle value oe
Answer: Incorrect: the median (19) is less than the lower quartile (22); the lower quartile can never be greater than the median.
Question 3
(a) B1 m ≤ 40 : 24 cao
(a) B1 m ≤ 80 : 68 cao
(a) Answer: 24 and 68
(b) B1 (60, 48) cao
(b) Answer: (60, 48)
Question 4
(a) M1 identifies the median position (40th value) and reads/interpolates
(a) A1 awrt 53 minutes (accept 51-56)
(a) Answer: awrt 53 minutes (accept 51-56)
(b) M1 identifies the lower quartile position (20th value) and reads/interpolates
(a) B1 recognises the lower boundary (-20.5) is negative
(a) B1 correct conclusion in context, e.g. screen time cannot be negative/less than 0 minutes, so no value could ever fall below the lower boundary oe
(a) Answer: The lower boundary is negative (-20.5), but screen time cannot be negative, so no value can be a low outlier.
(b) M1 compares 150 with the upper boundary (127.5)
(b) A1 ft: yes, 150 is an outlier since 150 > 127.5
(b) Answer: Yes, 150 minutes is an outlier, since 150 > 127.5.
Question 7
M1 reads/interpolates the cumulative frequency at m = 127.5 from the graph
A1 cumulative frequency awrt 78 (76-80)
B1 ft: number of students = awrt 2 (accept 1-3)
Answer: Approximately 2 students (accept 1-3)
Question 8
B1 ft: box drawn from the lower quartile (35) to the upper quartile (72), with the median (53) marked inside the box
B1 lower whisker drawn from the box to the minimum, non-outlier value (3)
B1 ft: upper whisker drawn from the box only as far as 118 (the largest non-outlier value), not as far as 150
B1 ft: the outlier at 150 marked as a separate point beyond the upper whisker, with a suitable linear scale and labelled axis
Answer: Box from 35 to 72 with median at 53; whiskers to 3 and 118; outlier plotted separately at 150.
Question 9
(a) M1 40/400 (=0.1) used as the sampling fraction, oe
(a) A1 18 boys and 22 girls cao
(a) Answer: 18 boys, 22 girls
(b) B1 valid advantage, e.g. it guarantees both boys and girls are represented in the sample in proportion to their numbers in the year group, making a comparison between them fairer/more reliable oe
(b) Answer: It ensures boys and girls are represented in proportion to the year group, making comparisons between them fairer.
(c) B1 recognises the number of hours is continuous/could take almost any value, so a list of 40 individual results would be hard to summarise/organise oe
(c) B1 recognises a cumulative frequency graph is built from running totals of grouped class frequencies, not raw individual values, so the data must be grouped into class intervals first oe
(c) Answer: Grouping organises the continuous raw data into class intervals, which is what a cumulative frequency graph is built from (running totals of grouped frequencies).
(d) B1 valid check, e.g. compare a known characteristic of the sample (such as the gender split or average age) with the same characteristic for the whole of Year 10 to see if they are similar oe
(d) Answer: Compare a known characteristic of the sample (e.g. the gender split) with the whole year group to check they match.
Question 10
(a) B1 27 cao
(a) Answer: 27 minutes
(b) M1 upper boundary = 55 + 1.5 x 27 (=95.5), oe
(b) A1 no, since 90 < 95.5, cao
(b) Answer: No, 90 is not an outlier, since 90 < 95.5
Question 11
(a) B1 Year 11's median (53 minutes) is higher than Year 8's (42 minutes), so Year 11 students typically spent longer on their phones oe
(a) Answer: Year 11 had the higher median (53 vs 42 minutes), so typically spent longer on their phones.
(b) B1 ft: Year 11's IQR (37 minutes) is greater than Year 8's (27 minutes), so screen times were more variable/spread out among Year 11 students oe
(b) Answer: Year 11's IQR (37) is greater than Year 8's (27), so Year 11 screen times were more variable.
(c) M1 correctly references the presence/absence of an outlier in each group
(c) A1 valid contextual comment, e.g. while most Year 8 students had broadly similar screen-time habits, at least one Year 11 student had an exceptionally high screen time compared with the rest of their year group, which was not seen in the Year 8 sample oe
(c) Answer: This suggests at least one Year 11 student's screen time was exceptionally high compared with the rest of their year group, unlike the more consistent Year 8 sample.
Question 12
B1 valid skew judgement with correct comparison, e.g. the lower whisker (3 to 35, a gap of 32) is shorter than the upper whisker (72 to 118, a gap of 46), so the distribution is positively (right) skewed oe
B1 recognises the outlier at 150 reinforces/increases this positive skew, extending the long tail even further to the right
B1 correct comparison of mean and median, e.g. the outlier would pull the mean up/increase it, since it is an extreme high value, but would have little or no effect on the median, so the mean would end up higher than the median oe
Answer: Positively (right) skewed; the outlier increases the skew and would inflate the mean but not the median.
Question 13
(a) M1 forms a correct equation, e.g. q + 1.5(q - 20) = 50, where q is the upper quartile
(c) M1 ft: compares each of the 5 masses with both boundaries
(c) A1 ft: 12 g (below 13.5, a low outlier) and 44 g (above 41.5, a high outlier); 22 g, 28 g and 33 g are not outliers
(c) Answer: 12 g and 44 g are outliers; 22 g, 28 g and 33 g are not.
Question 15
(a) B1 median and interquartile range cao
(a) B1 valid reason, e.g. the mean and standard deviation are strongly affected/distorted by extreme values, whereas the median and IQR are resistant to (not affected by) outliers oe
(a) Answer: Median and interquartile range, because they are not distorted by the two outliers, unlike the mean and standard deviation.
(b) B1 the mean would decrease/drop noticeably, since it currently includes two very large values which inflate the total oe
(b) B1 the median would change little or not at all, since it depends only on the middle value(s) of the ordered data, not the extreme values oe
(b) Answer: The mean would drop noticeably, but the median would stay about the same.
Question 16
(a) M1 identifies the median position (60th value) and the lower/upper quartile positions (30th and 90th values) and reads/interpolates
(a) M1 ft: compares 30 and 118 with both boundaries (awrt 35 and awrt 103)
(a) A1 ft: both are outliers, since 30 < 35 (low outlier) and 118 > 103 (high outlier)
(a) Answer: Both are outliers: 30 bpm is a low outlier and 118 bpm is a high outlier.
(b) B1 ft: box drawn from the lower quartile (awrt 60.5) to the upper quartile (77.5), with the median (awrt 68) marked inside the box
(b) B1 whiskers drawn correctly from the box to the minimum (44) and the maximum (96) of the routine sample
(b) B1 correct linear scale used and axis labelled; no outliers marked, since 44 and 96 both lie within the boundaries
(b) Answer: Box from awrt 60.5 to 77.5 with median at awrt 68; whiskers to 44 and 96; no outliers.
(c) B1 recognises the two readings were collected under very different circumstances (a medical emergency, not a routine check), so may not represent the same population as the 120 routine patients
(c) B1 concludes it would not be appropriate to combine them without care, since doing so could distort the analysis of what counts as a 'normal' resting heart rate oe; links to checking data validity during the processing/interpreting stages of the statistical enquiry cycle
(c) Answer: No, not without care; the two emergency readings come from a different context/population, and including them could distort the analysis of 'normal' resting heart rates.