A cycling club logs the distance, in km, ridden by each of its 8 members on a Saturday club run. The office assistant's raw log is shown below.
Row
Rider
Distance recorded
1
Priya
42 km
2
Tom
38 km
3
Freya
45 km
4
Oscar
51 km
5
Freya
45 km
6
Nadia
29 miles
7
Callum
47 km
8
Isla
33 km
(a)State which two row numbers are a duplicate entry for the same ride.(1)
(b)Explain why recording row 6's distance in miles, when every other row is in km, could cause a problem when the club works out the total distance ridden.(1)
(c)Given that 1 mile = 1.6 km, convert Nadia's distance to km, correct to 1 decimal place.(2)
(Total for Question 1 is 4 marks)
2
A small museum records the number of visitors on each of 11 consecutive days.
Day
1
2
3
4
5
6
7
8
9
10
11
Visitors
84
79
91
88
76
95
810
82
89
77
85
(a)Write down the value in the list that is most likely to be a data-entry error.(1)
(b)Suggest what the correct number of visitors was likely to be on day 7, and explain your reasoning.(1)
(c)Work out the range of the visitor numbers exactly as recorded, including the value 810.(2)
(d)The value is corrected to 81. Work out the range of the corrected data set.(2)
(e)Explain the effect that this data-entry error had on the range.(1)
(Total for Question 2 is 7 marks)
3
For each statement about data cleaning and outliers, write down whether it is True or False.
(i)The rule Q1 - 1.5 x IQR is used to find an upper boundary for identifying outliers.(1)
(ii)A single very large outlier typically increases the standard deviation of a data set.(1)
(iii)If a missing value is imputed using the mean of the remaining data, the overall mean of the completed data set stays the same.(1)
(iv)Deleting a data point simply because it looks unusual, without first checking whether it is a genuine value, is good statistical practice.(1)
(Total for Question 3 is 4 marks)
4
A fitness app records the number of steps walked by 13 participants in a charity step challenge on one day, listed here in ascending order.
(b)Find the lower quartile (Q1) and the upper quartile (Q3) of the data.(2)
(c)Work out the interquartile range (IQR).(1)
(Total for Question 4 is 4 marks)
5
The step-count data from Question 4 is used again here: Q1 = 4750, Q3 = 6000 and IQR = 1250.
(a)A value is considered an outlier if it lies below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Work out these two boundaries for the step-count data.(2)
(b)Use your boundaries to identify any outlier(s) in the step-count data set.(1)
(c)The participant with 14200 steps says they took part in a long charity walk that day, in addition to their normal daily activity. Explain whether this participant's data should be deleted from the step-count data set.(2)
(Total for Question 5 is 5 marks)
6
A cafe records the number of customers served in the first 15 minutes of each of 8 hours during a shift.
Hour
1
2
3
4
5
6
7
8
Customers served
22
25
19
24
21
23
20
26
(a)Work out the mean number of customers served.(2)
(b)Using the formula standard deviation = square root of [ (sum of (x - mean)2) / n ], work out the standard deviation of the number of customers served. Give your answer correct to 2 decimal places.(3)
(Total for Question 6 is 5 marks)
7
Later that day, the same cafe runs a one-hour promotion. In this extra hour, 74 customers were served in the first 15 minutes, far more than in any of the 8 hours in Question 6. Including this promotional hour, the full data set of 9 values is:
19, 20, 21, 22, 23, 24, 25, 26, 74
(a)Work out the new mean number of customers, including the promotional hour. Give your answer correct to 2 decimal places.(2)
(b)Write down the new median.(1)
(c)Work out the new standard deviation of the 9 values, using the same formula as in Question 6. Give your answer correct to 2 decimal places.(3)
(d)By comparing your answers here with the mean (22.5), median (22.5) and standard deviation (2.29) from Question 6, comment on which measure changed the most because of the promotional hour, and which changed the least.(2)
(Total for Question 7 is 8 marks)
8
The box plot shows the time, in minutes, taken by 40 runners to finish a fun run. One runner's time is marked separately with a dot, since it was not included when the box and whiskers were drawn.
(a)Write down the median time and the interquartile range (IQR) shown on the box plot.(2)
(b)One runner's time is marked separately at 71 minutes. Using the rule that a value is an outlier if it is above Q3 + 1.5 x IQR, show that 71 minutes is an outlier.(2)
(c)Give a possible genuine reason why this runner's time was so much higher than everyone else's, rather than assuming it must be a data-entry error.(1)
(Total for Question 8 is 5 marks)
9
A gardener records the height, in cm, of 15 sunflower plants at the end of the growing season, listed here in ascending order.
(a)Find the median, the lower quartile (Q1) and the upper quartile (Q3) of the 15 heights.(3)
(b)Show that the tallest sunflower (210 cm) is an outlier, using the boundary Q3 + 1.5 x IQR.(2)
(c)On the grid provided, draw a box plot for the 14 non-outlier heights (that is, excluding the 210 cm sunflower). Plot the 210 cm sunflower separately, as an isolated point beyond the whisker.(3)
(Total for Question 9 is 8 marks)
10
A delivery company records the time, in minutes, taken for 10 parcels to be delivered from the depot to customers' homes.
Parcel
1
2
3
4
5
6
7
8
9
10
Time (min)
18
22
(not recorded)
19
145
21
20
17
23
19
(a)Parcel 3's delivery time was not recorded. Give one advantage and one disadvantage of imputing (estimating) this missing value using the mean of the other delivery times, rather than deleting parcel 3's record entirely.(2)
(b)Parcel 5's delivery time is recorded as 145 minutes, far higher than every other parcel. Explain why the company should investigate this value rather than immediately deleting it or replacing it with the mean.(2)
(c)The company decides to impute parcel 3's missing value using the mean of the other 9 recorded times (treating parcel 5's 145 minutes as genuine for now). Work out this imputed value, correct to 1 decimal place.(2)
(d)Comment on whether this imputed value from part (c) is a sensible estimate for parcel 3's delivery time.(1)
(Total for Question 10 is 7 marks)
11
Aaron is investigating how many hours Year 11 students spend revising per week. He designs a questionnaire, distributes it to a random sample of 30 students, and enters their responses into a spreadsheet. (The statistical enquiry cycle has four stages: plan, collect, process, interpret.)
(a)Aaron notices that two students have left the revising-hours question blank. At which stage of the statistical enquiry cycle should Aaron decide how to deal with these missing responses?(1)
(b)One response reads '20 hours', much higher than every other response (all between 2 and 12 hours). Aaron calculates that Q1 = 4 hours and Q3 = 9 hours for the sample. Use the rule Q3 + 1.5 x IQR to decide whether 20 hours should be treated as an outlier.(3)
(c)Aaron decides to leave the two blank responses out of his analysis rather than guessing a value for them, but to investigate the '20 hours' response further rather than deleting it immediately. Explain why both of these are reasonable decisions.(2)
(Total for Question 11 is 6 marks)
12
A quality-control inspector at a factory measures the mass, in grams, of 10 bags of rice from a production line.
Bag
1
2
3
4
5
6
7
8
9
10
Mass (g)
501
498
503
50
502
500
499
497
504
499
(a)Identify the bag whose mass is very likely a data-entry error.(1)
(b)State the value this entry should probably be corrected to, and explain your reasoning.(1)
(c)Using the RAW data (that is, with 50 g for bag 4), work out the mean and standard deviation of the 10 masses. Give the standard deviation correct to 1 decimal place. You may use the formula standard deviation = square root of [ (sum of x2)/n - mean2 ].(4)
(d)Using the CORRECTED data (that is, with 500 g for bag 4), work out the mean and standard deviation of the 10 masses. Give the standard deviation correct to 1 decimal place.(3)
(e)Comment on the effect that this single data-entry error had on the standard deviation.(1)
(Total for Question 12 is 10 marks)
13
A machine operator records the diameter, in mm, of 15 ball bearings and calculates two summary totals from her spreadsheet: the sum of the diameters, Σx = 309.6, and the sum of the squared diameters, Σx² = 6476.86.
(a)Using Σx = 309.6, work out the mean diameter as originally recorded.(1)
(b)A check of the machine's calibration log shows that one ball bearing's diameter was recorded as 29.6 mm, but should have been 19.6 mm (a digit was mistyped). Work out the corrected value of Σx.(2)
(c)Work out the corrected value of Σx².(2)
(d)Hence work out the corrected mean and standard deviation of the 15 diameters. Give the standard deviation correct to 2 decimal places. You may use the formula standard deviation = square root of [ (sum of x2)/n - mean2 ].(3)
(e)Given that the standard deviation calculated from the original (uncorrected) Σx and Σx² was 2.40 mm (2 dp), comment on the effect that correcting the single misrecorded value had on the standard deviation.(1)
(Total for Question 13 is 9 marks)
14
A car park attendant's spreadsheet automatically logs the length of stay, in minutes, of vehicles using the car park. While checking the log for one day, she discovers that 6 vehicles which left during the 60 < t ≤ 90 minute class were also mistakenly logged a second time in the 90 < t ≤ 120 minute class, because of a fault in the automatic timer. The table below shows the frequencies as they were first recorded (before this fault was found).
Length of stay, t (min)
0 < t ≤ 30
30 < t ≤ 60
60 < t ≤ 90
90 < t ≤ 120
Frequency (as recorded)
8
14
23
21
(a)Explain why the total of the frequencies in the table (66) does not match the 60 vehicles that actually used the car park that day.(1)
(b)Write down the corrected frequency for the 90 < t ≤ 120 class.(1)
(c)Using mid-interval values, work out an estimate for the mean length of stay using the corrected frequencies (8, 14, 23, 15).(3)
(d)Given that Σfx² = 324900 for the corrected data, work out an estimate for the standard deviation of the corrected data. Give your answer correct to 1 decimal place.(2)
(e)Using the UNCORRECTED (originally recorded) frequencies, the estimated mean is 70.9 minutes and the estimated standard deviation is 29.9 minutes (both 1 dp). Comment on the effect that the duplicate-logging error had on these estimates.(1)
(Total for Question 14 is 8 marks)
15
The table below shows three measures of spread for the rice-bag masses from Question 12, calculated before and after the data-entry error was corrected.
Measure of spread
Before cleaning (raw)
After cleaning (corrected)
Range
454 g
7 g
Interquartile range (IQR)
4 g
3 g
Standard deviation
135.1 g
2.1 g
(a)By comparing the 'before' and 'after' values in the table, state which measure of spread changed the LEAST as a result of cleaning the data.(1)
(b)Explain why the interquartile range is much less affected than the range or the standard deviation by a single extreme/erroneous value.(2)
(c)A different data set is known to contain at least one extreme data-entry error, but there is no time to check and correct every entry before a report is due. Recommend which of the three measures of spread in the table the analyst should use to describe the data, and justify your recommendation.(2)
(Total for Question 15 is 5 marks)
Mark scheme · H01 Data Cleaning and Outliers
Question 1
(a) B1 3 and 5 cao
(a) Answer: Rows 3 and 5
(b) B1 explains that mixing units would make the total/sum of distances meaningless unless all values are converted to the same unit first oe
(b) Answer: If the 29 is added to the other distances without converting it, the total distance ridden would be wrong, because 29 miles is not the same as 29 km.
(c) M1 29 x 1.6, oe
(c) A1 46.4 km cao
(c) Answer: 46.4 km
Question 2
(a) B1 810 cao
(a) Answer: 810
(b) B1 81 (an extra 0 was probably typed after 81), oe reasoning that links to the range of the other values
(b) Answer: 81 visitors (an extra 0 was probably typed by mistake).
(c) M1 810 - 76, oe
(c) A1 734 cao
(c) Answer: 734
(d) M1 95 - 76, oe
(d) A1 19 cao
(d) Answer: 19
(e) B1 the error made the range far larger than it should be (734 instead of 19), because the range only uses the two most extreme values, so one incorrect extreme value distorts it heavily oe
(e) Answer: The error inflated the range hugely (734 instead of the true 19), because the range depends only on the biggest and smallest values, so a single wrong extreme value badly distorts it.
(c) B1 the value is a genuine result, explained by real behaviour, rather than a data-entry error oe
(c) B1 it should be kept (and possibly flagged/analysed separately) rather than deleted, since deleting genuine data would lose real information and could bias the results oe
(c) Answer: No, it should not be deleted; 14200 is a genuine value explained by the charity walk, so removing it would lose real information and could bias the results.
Question 6
(a) M1 22+25+19+24+21+23+20+26, oe (=180)
(a) A1 22.5 cao
(a) Answer: 22.5
(b) M1 finds the deviation of each value from the mean (22.5) and squares each one
(b) M1 sums the squared deviations and divides by n=8, oe (=5.25)
(b) A1 2.29 (awrt 2.29) cao
(b) Answer: 2.29 (awrt)
Question 7
(a) M1 19+20+21+22+23+24+25+26+74, oe (=254)
(a) A1 28.22 (awrt 28.22) cao
(a) Answer: 28.22 (awrt)
(b) B1 23 cao
(b) Answer: 23
(c) M1 finds the deviation of each value from the new mean and squares each one
(c) M1 sums the squared deviations and divides by n=9, oe
(c) A1 16.33 (awrt 16.33) cao
(c) Answer: 16.33 (awrt)
(d) B1 the standard deviation changed by far the most, rising from 2.29 to about 16.33, roughly 7 times bigger oe
(d) B1 the median changed the least, rising only slightly from 22.5 to 23, while the mean rose more noticeably to 28.22, showing the median is the measure most resistant to an extreme value oe
(d) Answer: The standard deviation changed the most (2.29 to 16.33); the median changed the least (22.5 to 23); the mean also rose noticeably (22.5 to 28.22). This shows the median is the most resistant to an outlier, while the standard deviation is the most sensitive.
Question 8
(a) B1 median = 38 cao
(a) B1 IQR = 44 - 32 = 12 cao
(a) Answer: Median = 38 minutes, IQR = 12 minutes
(b) M1 44 + 1.5x12, oe
(b) A1 62, and 71 > 62 so it is an outlier, cao
(b) Answer: 62 minutes; since 71 > 62, it is an outlier.
(c) B1 e.g. the runner may have stopped to help an injured runner, walked most of the course, or been injured themselves, oe
(c) Answer: The runner may have stopped to help someone else, been injured, or walked most of the course, so 71 minutes could be a genuine time.
Question 9
(a) B1 median = 156 cao
(a) B1 Q1 = 150 cao
(a) B1 Q3 = 163 cao
(a) Answer: Median = 156 cm, Q1 = 150 cm, Q3 = 163 cm
(b) M1 IQR = 13 and 163 + 1.5x13, oe
(b) A1 182.5, and 210 > 182.5 so it is an outlier, cao
(b) Answer: 182.5 cm; since 210 > 182.5, it is an outlier.
(c) B1 box drawn correctly from Q1 = 150 to Q3 = 163 with the median line at 156
(c) B1 whiskers drawn correctly to the minimum (142) and to the maximum of the remaining values (168)
(c) B1 the 210 cm sunflower plotted as a separate point, clearly apart from the whisker
(c) Answer: Box from 150 to 163 with median line at 156; whiskers to 142 and 168; the 210 cm sunflower plotted as a separate point.
Question 10
(a) B1 advantage: keeps the sample size at 10 / keeps some information about parcel 3, so later calculations can still use all the deliveries oe
(a) B1 disadvantage: the imputed value is not a genuine measurement and may not reflect parcel 3's true delivery time, especially if it was unusual oe
(a) Answer: Advantage: keeps the full sample of 10 parcels for later calculations. Disadvantage: the imputed value is not a real measurement and may not reflect what actually happened to parcel 3.
(b) B1 145 minutes might reveal a genuine problem, e.g. a lost or delayed parcel, which is valuable information for the business oe
(b) B1 deleting or replacing it straight away would hide this problem and could make future delivery-time estimates unrealistically low/optimistic oe
(b) Answer: 145 minutes could reveal a genuine delivery problem (e.g. a lost parcel); deleting or replacing it immediately would hide this and could make future delivery-time estimates too optimistic.
(c) M1 18+22+19+145+21+20+17+23+19, oe (=304), divided by 9
(c) A1 33.8 (awrt 33.8) cao
(c) Answer: 33.8 minutes (awrt)
(d) B1 no - because parcel 5's extreme 145 minutes was included in the calculation, it pulls the mean far higher than a typical delivery (all other times are between 17 and 23 minutes), so 33.8 minutes is not realistic; parcel 5 should be resolved before imputing parcel 3 oe
(d) Answer: No; including the extreme 145-minute delivery pulls the imputed value up to 33.8 minutes, far higher than a typical delivery time (17 to 23 minutes), so it is not a sensible estimate.
Question 11
(a) B1 process cao
(a) Answer: The process stage.
(b) M1 IQR = 9 - 4 = 5, oe
(b) M1 9 + 1.5x5, oe
(b) A1 16.5, and 20 > 16.5 so yes it is an outlier, cao
(b) Answer: Yes; the boundary is 16.5 hours, and 20 hours is greater than this.
(c) B1 blank responses contain no genuine information, so guessing a value could introduce bias/inaccuracy that a real value would not have oe
(c) B1 an unusually high but not impossible value like 20 hours could be genuine, e.g. a student revising intensively before exams, so deleting it straight away could lose real, useful information oe
(c) Answer: Blank responses have no genuine data to guess from, so estimating them risks bias. 20 hours could be a genuine value (e.g. a student cramming before exams), so it should be checked rather than deleted straight away.
Question 12
(a) B1 bag 4 (50 g) cao
(a) Answer: Bag 4 (50 g).
(b) B1 500 g - a zero has probably been dropped when the mass was typed in, oe
(b) Answer: 500 g (a zero was probably dropped when it was typed in).
(c) M1 sum = 4553, oe
(c) A1 mean = 455.3 cao
(c) M1 correct substitution into the formula, e.g. 2255545/10 - 455.32, oe
(c) A1 135.1 (awrt 135.1) cao
(c) Answer: Mean = 455.3 g, standard deviation = 135.1 g (awrt).
(d) M1 sum = 5003, oe
(d) A1 mean = 500.3 cao
(d) A1 SD = 2.1 cao
(d) Answer: Mean = 500.3 g, standard deviation = 2.1 g.
(e) B1 the error inflated the standard deviation enormously, from 2.1 g to about 135.1 g (more than 60 times larger), showing that even one incorrect value can make the spread look far bigger than it really is oe
(e) Answer: The error inflated the standard deviation hugely, from 2.1 g (corrected) to about 135.1 g (raw), more than 60 times larger, showing how much one wrong value can distort a measure of spread.
Question 13
(a) B1 20.64 mm cao
(a) Answer: 20.64 mm.
(b) M1 309.6 - 29.6 + 19.6, oe
(b) A1 299.6 cao
(b) Answer: Σx = 299.6.
(c) M1 6476.86 - 29.62 + 19.62, oe
(c) A1 5984.86 cao
(c) Answer: Σx² = 5984.86.
(d) M1 mean = 299.6/15, oe
(d) A1 19.97 (awrt 19.97) cao
(d) A1 SD = 0.24 (awrt 0.24) cao
(d) Answer: Mean = 19.97 mm (awrt), standard deviation = 0.24 mm (awrt).
(e) B1 correcting the value reduced the standard deviation hugely, from about 2.40 mm to about 0.24 mm (roughly one tenth of its original size), showing how much a single incorrect entry can inflate a standard deviation oe
(e) Answer: The standard deviation fell dramatically, from about 2.40 mm (using the recorded value) to about 0.24 mm (corrected), about one tenth the size, showing how much one wrong entry can inflate the spread.
Question 14
(a) B1 the 6 vehicles that left during the 60-90 min class were logged twice - once correctly and once in error in the 90-120 min class - so these 6 extra rows do not represent extra vehicles oe
(a) Answer: The 6 vehicles that left in the 60-90 minute class were logged a second time by mistake in the 90-120 minute class, so those 6 extra entries do not represent extra vehicles.
(b) B1 15 cao (21 - 6)
(b) Answer: 15.
(c) M1 mid-values 15, 45, 75, 105 used, oe
(c) M1 sum of f x mid-value = 4050, oe
(c) A1 mean = 67.5 minutes cao
(c) Answer: 67.5 minutes.
(d) M1 324900/60 - 67.52, oe
(d) A1 29.3 (awrt 29.3) cao
(d) Answer: 29.3 minutes (awrt).
(e) B1 the error increased both the estimated mean (by about 3.4 minutes) and the estimated standard deviation (by about 0.6 minutes), because it wrongly added extra vehicles into a higher time class oe
(e) Answer: The error increased the estimated mean (by about 3.4 minutes) and the estimated standard deviation (by about 0.6 minutes), since it wrongly shifted 6 extra 'vehicles' into a higher time class than they really belonged to.
Question 15
(a) B1 the interquartile range (IQR) cao - changed from 4 g to 3 g, a change of only 1 g
(a) Answer: The interquartile range (IQR), which changed from 4 g to only 3 g.
(b) B1 the IQR only uses the middle 50% of the data (the values between Q1 and Q3), so a single extreme value at either end has no direct effect on it oe
(b) B1 the range uses only the two most extreme values, and the standard deviation uses every squared deviation from the mean, so both are heavily influenced by a single very large or very small value oe
(b) Answer: The IQR only depends on the middle 50% of the data, so an extreme value at either end does not affect it directly. The range depends only on the two most extreme values, and the standard deviation depends on every squared deviation from the mean, so both are strongly affected by a single extreme or erroneous value.
(c) B1 recommends the interquartile range (IQR)
(c) B1 justifies this because the IQR is more resistant/robust to outliers and errors than the range or standard deviation, so it will still give a fair description of the spread of the majority of the data even if some values are wrong oe
(c) Answer: The analyst should use the interquartile range (IQR), because it is far more resistant to outliers and data-entry errors than the range or standard deviation, so it will still describe the spread of the data fairly even with an uncorrected error present.