Data Cleaning and Outliers - Worksheets, Questions and Revision

15 original exam-style questions - 13 pages of questions with a full mark scheme - free printable PDF.

Download PDFJump to mark scheme (page 14)
« Previous: Recognising Misleading GraphsNext: Correlation and Causation »
Revision Library
revisionlibrary.co.uk
HIGHER

H01 Data Cleaning and Outliers

EDEXCEL 1ST0 · Calculator allowed · about 90 minutes
Total Marks
Name: _______________________________    Date: ____ / ____ / ______
Answer ALL questions. Show all your working.
1
A cycling club logs the distance, in km, ridden by each of its 8 members on a Saturday club run. The office assistant's raw log is shown below.
RowRiderDistance recorded
1Priya42 km
2Tom38 km
3Freya45 km
4Oscar51 km
5Freya45 km
6Nadia29 miles
7Callum47 km
8Isla33 km
(a)State which two row numbers are a duplicate entry for the same ride.(1)
(b)Explain why recording row 6's distance in miles, when every other row is in km, could cause a problem when the club works out the total distance ridden.(1)
(c)Given that 1 mile = 1.6 km, convert Nadia's distance to km, correct to 1 decimal place.(2)
(Total for Question 1 is 4 marks)
2
A small museum records the number of visitors on each of 11 consecutive days.
Day1234567891011
Visitors84799188769581082897785
(a)Write down the value in the list that is most likely to be a data-entry error.(1)
(b)Suggest what the correct number of visitors was likely to be on day 7, and explain your reasoning.(1)
(c)Work out the range of the visitor numbers exactly as recorded, including the value 810.(2)
(d)The value is corrected to 81. Work out the range of the corrected data set.(2)
(e)Explain the effect that this data-entry error had on the range.(1)
(Total for Question 2 is 7 marks)
3
For each statement about data cleaning and outliers, write down whether it is True or False.
(i)The rule Q1 - 1.5 x IQR is used to find an upper boundary for identifying outliers.(1)
(ii)A single very large outlier typically increases the standard deviation of a data set.(1)
(iii)If a missing value is imputed using the mean of the remaining data, the overall mean of the completed data set stays the same.(1)
(iv)Deleting a data point simply because it looks unusual, without first checking whether it is a genuine value, is good statistical practice.(1)
(Total for Question 3 is 4 marks)
4
A fitness app records the number of steps walked by 13 participants in a charity step challenge on one day, listed here in ascending order.

4200, 4550, 4700, 4800, 5100, 5300, 5450, 5600, 5750, 5900, 6100, 6300, 14200
(a)Write down the median number of steps.(1)
(b)Find the lower quartile (Q1) and the upper quartile (Q3) of the data.(2)
(c)Work out the interquartile range (IQR).(1)
(Total for Question 4 is 4 marks)
5
The step-count data from Question 4 is used again here: Q1 = 4750, Q3 = 6000 and IQR = 1250.
(a)A value is considered an outlier if it lies below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Work out these two boundaries for the step-count data.(2)
(b)Use your boundaries to identify any outlier(s) in the step-count data set.(1)
(c)The participant with 14200 steps says they took part in a long charity walk that day, in addition to their normal daily activity. Explain whether this participant's data should be deleted from the step-count data set.(2)
(Total for Question 5 is 5 marks)
6
A cafe records the number of customers served in the first 15 minutes of each of 8 hours during a shift.
Hour12345678
Customers served2225192421232026
(a)Work out the mean number of customers served.(2)
(b)Using the formula standard deviation = square root of [ (sum of (x - mean)2) / n ], work out the standard deviation of the number of customers served. Give your answer correct to 2 decimal places.(3)
(Total for Question 6 is 5 marks)
7
Later that day, the same cafe runs a one-hour promotion. In this extra hour, 74 customers were served in the first 15 minutes, far more than in any of the 8 hours in Question 6. Including this promotional hour, the full data set of 9 values is:

19, 20, 21, 22, 23, 24, 25, 26, 74
(a)Work out the new mean number of customers, including the promotional hour. Give your answer correct to 2 decimal places.(2)
(b)Write down the new median.(1)
(c)Work out the new standard deviation of the 9 values, using the same formula as in Question 6. Give your answer correct to 2 decimal places.(3)
(d)By comparing your answers here with the mean (22.5), median (22.5) and standard deviation (2.29) from Question 6, comment on which measure changed the most because of the promotional hour, and which changed the least.(2)
(Total for Question 7 is 8 marks)
8
The box plot shows the time, in minutes, taken by 40 runners to finish a fun run. One runner's time is marked separately with a dot, since it was not included when the box and whiskers were drawn.
20 30 40 50 60 70 80 Time (minutes) 71
(a)Write down the median time and the interquartile range (IQR) shown on the box plot.(2)
(b)One runner's time is marked separately at 71 minutes. Using the rule that a value is an outlier if it is above Q3 + 1.5 x IQR, show that 71 minutes is an outlier.(2)
(c)Give a possible genuine reason why this runner's time was so much higher than everyone else's, rather than assuming it must be a data-entry error.(1)
(Total for Question 8 is 5 marks)
9
A gardener records the height, in cm, of 15 sunflower plants at the end of the growing season, listed here in ascending order.

142, 145, 148, 150, 151, 153, 154, 156, 158, 159, 161, 163, 165, 168, 210
130 140 150 160 170 180 190 200 210 220 Height (cm)
(a)Find the median, the lower quartile (Q1) and the upper quartile (Q3) of the 15 heights.(3)
(b)Show that the tallest sunflower (210 cm) is an outlier, using the boundary Q3 + 1.5 x IQR.(2)
(c)On the grid provided, draw a box plot for the 14 non-outlier heights (that is, excluding the 210 cm sunflower). Plot the 210 cm sunflower separately, as an isolated point beyond the whisker.(3)
(Total for Question 9 is 8 marks)
10
A delivery company records the time, in minutes, taken for 10 parcels to be delivered from the depot to customers' homes.
Parcel12345678910
Time (min)1822(not recorded)191452120172319
(a)Parcel 3's delivery time was not recorded. Give one advantage and one disadvantage of imputing (estimating) this missing value using the mean of the other delivery times, rather than deleting parcel 3's record entirely.(2)
(b)Parcel 5's delivery time is recorded as 145 minutes, far higher than every other parcel. Explain why the company should investigate this value rather than immediately deleting it or replacing it with the mean.(2)
(c)The company decides to impute parcel 3's missing value using the mean of the other 9 recorded times (treating parcel 5's 145 minutes as genuine for now). Work out this imputed value, correct to 1 decimal place.(2)
(d)Comment on whether this imputed value from part (c) is a sensible estimate for parcel 3's delivery time.(1)
(Total for Question 10 is 7 marks)
11
Aaron is investigating how many hours Year 11 students spend revising per week. He designs a questionnaire, distributes it to a random sample of 30 students, and enters their responses into a spreadsheet. (The statistical enquiry cycle has four stages: plan, collect, process, interpret.)
(a)Aaron notices that two students have left the revising-hours question blank. At which stage of the statistical enquiry cycle should Aaron decide how to deal with these missing responses?(1)
(b)One response reads '20 hours', much higher than every other response (all between 2 and 12 hours). Aaron calculates that Q1 = 4 hours and Q3 = 9 hours for the sample. Use the rule Q3 + 1.5 x IQR to decide whether 20 hours should be treated as an outlier.(3)
(c)Aaron decides to leave the two blank responses out of his analysis rather than guessing a value for them, but to investigate the '20 hours' response further rather than deleting it immediately. Explain why both of these are reasonable decisions.(2)
(Total for Question 11 is 6 marks)
12
A quality-control inspector at a factory measures the mass, in grams, of 10 bags of rice from a production line.
Bag12345678910
Mass (g)50149850350502500499497504499
(a)Identify the bag whose mass is very likely a data-entry error.(1)
(b)State the value this entry should probably be corrected to, and explain your reasoning.(1)
(c)Using the RAW data (that is, with 50 g for bag 4), work out the mean and standard deviation of the 10 masses. Give the standard deviation correct to 1 decimal place. You may use the formula standard deviation = square root of [ (sum of x2)/n - mean2 ].(4)
(d)Using the CORRECTED data (that is, with 500 g for bag 4), work out the mean and standard deviation of the 10 masses. Give the standard deviation correct to 1 decimal place.(3)
(e)Comment on the effect that this single data-entry error had on the standard deviation.(1)
(Total for Question 12 is 10 marks)
13
A machine operator records the diameter, in mm, of 15 ball bearings and calculates two summary totals from her spreadsheet: the sum of the diameters, Σx = 309.6, and the sum of the squared diameters, Σx² = 6476.86.
(a)Using Σx = 309.6, work out the mean diameter as originally recorded.(1)
(b)A check of the machine's calibration log shows that one ball bearing's diameter was recorded as 29.6 mm, but should have been 19.6 mm (a digit was mistyped). Work out the corrected value of Σx.(2)
(c)Work out the corrected value of Σx².(2)
(d)Hence work out the corrected mean and standard deviation of the 15 diameters. Give the standard deviation correct to 2 decimal places. You may use the formula standard deviation = square root of [ (sum of x2)/n - mean2 ].(3)
(e)Given that the standard deviation calculated from the original (uncorrected) Σx and Σx² was 2.40 mm (2 dp), comment on the effect that correcting the single misrecorded value had on the standard deviation.(1)
(Total for Question 13 is 9 marks)
14
A car park attendant's spreadsheet automatically logs the length of stay, in minutes, of vehicles using the car park. While checking the log for one day, she discovers that 6 vehicles which left during the 60 < t ≤ 90 minute class were also mistakenly logged a second time in the 90 < t ≤ 120 minute class, because of a fault in the automatic timer. The table below shows the frequencies as they were first recorded (before this fault was found).
Length of stay, t (min)0 < t ≤ 3030 < t ≤ 6060 < t ≤ 9090 < t ≤ 120
Frequency (as recorded)8142321
(a)Explain why the total of the frequencies in the table (66) does not match the 60 vehicles that actually used the car park that day.(1)
(b)Write down the corrected frequency for the 90 < t ≤ 120 class.(1)
(c)Using mid-interval values, work out an estimate for the mean length of stay using the corrected frequencies (8, 14, 23, 15).(3)
(d)Given that Σfx² = 324900 for the corrected data, work out an estimate for the standard deviation of the corrected data. Give your answer correct to 1 decimal place.(2)
(e)Using the UNCORRECTED (originally recorded) frequencies, the estimated mean is 70.9 minutes and the estimated standard deviation is 29.9 minutes (both 1 dp). Comment on the effect that the duplicate-logging error had on these estimates.(1)
(Total for Question 14 is 8 marks)
15
The table below shows three measures of spread for the rice-bag masses from Question 12, calculated before and after the data-entry error was corrected.
Measure of spreadBefore cleaning (raw)After cleaning (corrected)
Range454 g7 g
Interquartile range (IQR)4 g3 g
Standard deviation135.1 g2.1 g
(a)By comparing the 'before' and 'after' values in the table, state which measure of spread changed the LEAST as a result of cleaning the data.(1)
(b)Explain why the interquartile range is much less affected than the range or the standard deviation by a single extreme/erroneous value.(2)
(c)A different data set is known to contain at least one extreme data-entry error, but there is no time to check and correct every entry before a report is due. Recommend which of the three measures of spread in the table the analyst should use to describe the data, and justify your recommendation.(2)
(Total for Question 15 is 5 marks)
Mark scheme · H01 Data Cleaning and Outliers

Question 1

Question 2

Question 3

Question 4

Question 5

Question 6

Question 7

Question 8

Question 9

Question 10

Question 11

Question 12

Question 13

Question 14

Question 15