A school office assistant is entering sign-up details for a new coding club into a spreadsheet. The raw data entered is shown below.
Entry
Name
Age (years)
Height
1
Aliyah Khan
14
1.58 m
2
Ben Ward
(not recorded)
1.62 m
3
Chloe Bird
13
152 cm
4
Chloe Bird
13
152 cm
(a)State the problem with entry 2.(1)
(b)State the problem with entries 3 and 4.(1)
(c)Entry 3's height is recorded in a different unit (cm) from the rest of the table (m). Explain why this inconsistency could cause a problem when the data is analysed.(1)
(Total for Question 1 is 3 marks)
2
A leisure centre's membership sign-up log was entered by two different receptionists on the same day.
Row
Name entered
Date of birth
1
Mohammed Iqbal
03/05/2011
2
M. Iqbal
03/05/2011
3
Sara Owusu
21/11/2010
4
Leo Fischer
09/02/2012
(a)Write down the two row numbers that most likely refer to the same person.(1)
(b)Give two reasons why rows 1 and 2 are likely to refer to the same person, even though the names were not entered identically.(2)
(Total for Question 2 is 3 marks)
3
As part of a science experiment, a technician measured the height, in cm, of 12 seedlings. The raw data recorded is shown below.
(a)Give two reasons why this list of responses needs to be cleaned before it can be used to make a frequency table.(2)
(b)After cleaning the data, complete the frequency table above for the three colours.(2)
(Total for Question 8 is 4 marks)
9
The weekly pocket money, in £, of 9 pupils in a class is recorded below.
£5, £6, £5, £7, £6, £5, £6, £7, £45
(a)Work out the mean and the median weekly pocket money for all 9 pupils.(3)
(b)One pupil's pocket money, £45, is much higher than the rest. Excluding the £45, work out the new mean and new median for the remaining 8 pupils.(3)
(c)Comment on the effect that removing the £45 had on the mean compared with its effect on the median.(1)
(Total for Question 9 is 7 marks)
10
Jasmine is investigating whether pupils in her year group spend more time on homework on weekdays or at weekends. She plans a survey, collects responses using a paper questionnaire, then enters the data into a spreadsheet before working out any averages. (The statistical enquiry cycle has four stages: plan, collect, process, interpret.)
(a)Jasmine notices that some pupils have written their homework time in minutes and others in hours. At which stage of the enquiry cycle should Jasmine correct this problem, before she calculates any averages?(1)
(b)Explain why it is important to clean the data during the process stage before moving on to the interpret stage.(1)
(c)Jasmine also finds that 3 questionnaires have been returned completely blank. Give one appropriate action she could take with these blank responses at the process stage.(1)
(d)State one reason it would NOT be sensible for Jasmine to guess plausible homework times to fill in the 3 blank questionnaires.(1)
(Total for Question 10 is 4 marks)
11
A gym's sign-in sheet for one evening class lists 10 entries, taken automatically from members scanning their cards.
Row
Name
Member ID
1
Grace Oduya
M1042
2
K. Yilmaz
M1077
3
Grace Oduya
M1042
4
Daniel Cross
M1055
5
Kerem Yilmaz
M1077
6
Priti Shah
M1061
7
Daniel Cross
M1055
8
Amir Bello
M1089
9
Grace Oduya
M1042
10
Priti Shah
M1061
(a)Work out how many unique members actually attended the class.(2)
(b)Write down the Member ID that was recorded the greatest number of times, and state how many times it appears.(2)
(Total for Question 11 is 4 marks)
12
A cafe records the volume of coffee sold in take-away cups one morning. Some volumes were entered in millilitres and some in litres by mistake.
Cup
A
B
C
D
E
Volume recorded
350
0.4
250
0.3
400
(a)Explain why the volumes recorded for cups B and D are likely to have been entered in the wrong unit.(1)
(b)Convert the volumes for cups B and D into millilitres, so that all five volumes use consistent units.(2)
(Total for Question 12 is 3 marks)
13
A clinic logs patient appointment dates. Two different members of staff entered the dates in different formats.
Patient
Date entered
P
04/07/2026
Q
2026-07-11
R
09/07/2026
S
2026-07-03
(a)Write down which format has been used for patient Q's appointment date.(1)
(b)Explain why mixing date formats like this in one column could cause a problem when the appointments are sorted into date order.(1)
(c)Rewrite patient Q's date using the same day/month/year format as the other patients.(1)
(Total for Question 13 is 3 marks)
14
For each statement about data cleaning, write down whether it is True or False.
(i)An outlier is always caused by a mistake when the data was recorded.(1)
(ii)Removing a genuine extreme value from a data set just because it looks unusual can make the data set less representative of what really happened.(1)
(iii)If a data set contains duplicate entries for the same person, the true sample size may be smaller than the number of rows recorded.(1)
(Total for Question 14 is 3 marks)
15
The number of pints of milk delivered to 11 houses on a street one morning is recorded below, in ascending order.
2, 2, 3, 3, 4, 4, 4, 5, 5, 6, 15
(a)Write down the median number of pints delivered.(1)
(b)Find the lower quartile (Q1) and the upper quartile (Q3) of the data.(2)
(c)Work out the interquartile range (IQR).(1)
(d)A value is an outlier if it is lower than Q1 - 1.5 x IQR or higher than Q3 + 1.5 x IQR. Work out these two boundaries.(2)
(e)Use your boundaries to decide whether the delivery of 15 pints is an outlier. Give a reason.(1)
(f)Mark the lower and upper boundaries on the number line below, and circle any data value(s) that are outliers.(1)
(Total for Question 15 is 8 marks)
16
A researcher is investigating the resting heart rate, in beats per minute (bpm), of 9 volunteers, as part of a health survey. The raw data collected is shown below (this is the process stage of the statistical enquiry cycle).
Row
1
2
3
4
5
6
7
8
9
Heart rate (bpm)
68
71
(missing)
70
690
72
69
132
70
(a)Row 3 has a missing heart rate. Suggest one appropriate way the researcher could deal with this missing value at the process stage.(1)
(b)The heart rate for row 5 is recorded as 690 bpm. Explain why this is almost certainly a data-entry error, and suggest what the true value is likely to be.(2)
(c)The heart rate for row 8 is recorded as 132 bpm. Give one reason why this value might be a genuine reading rather than a data-entry error.(1)
(d)Explain why it would not be appropriate to treat rows 5 and 8 in the same way when cleaning this data set.(1)
(e)Row 3's true value is later found to be 70 bpm, and row 5 has been corrected to 69 bpm. Using these corrected values, work out the mean heart rate of all 9 volunteers, including the genuine value of 132 bpm.(2)
(f)Without volunteer 8's reading of 132 bpm, the mean of the remaining 8 volunteers would be 69.9 bpm (1 dp). Comment on the effect that including the genuine value of 132 bpm has on the mean.(1)
(Total for Question 16 is 8 marks)
Mark scheme · F01 Data Cleaning and Outliers
Question 1
(a) B1 identifies that the age has not been recorded / is missing oe
(a) Answer: Ben Ward's age is missing (not recorded).
(b) B1 identifies that entries 3 and 4 are duplicate entries for the same person oe
(b) Answer: Entries 3 and 4 are duplicates (the same person, Chloe Bird, has been entered twice).
(c) B1 explains that mixing units could lead to heights being compared or averaged incorrectly (e.g. 152 could be mistaken for 152 m) oe
(c) Answer: If the units are not made consistent, 152 could wrongly be treated as if it were in metres (152 m) when calculating an average or comparing heights, giving a meaningless result.
Question 2
(a) B1 1 and 2 cao
(a) Answer: Rows 1 and 2
(b) B1 both rows have the same date of birth (03/05/2011)
(b) B1 'M. Iqbal' is an abbreviated/shortened version of the name 'Mohammed Iqbal' oe
(b) Answer: Both rows share the same date of birth, and 'M. Iqbal' is just a shortened version of the full name 'Mohammed Iqbal'.
Question 3
(a) B1 1350 cao
(a) Answer: 1350
(b) B1 1350 cm (13.5 m) is an impossible height for a seedling, so it cannot be a genuine measurement oe
(b) B1 135 cm (an extra 0 was probably typed, or the decimal point misplaced) oe
(b) Answer: 1350 cm is impossible for a seedling, so it is a data-entry error; the true value was probably 135 cm.
Question 4
(a) M1 1350 - 129 oe
(a) A1 1221 cao
(a) Answer: 1221 cm
(b) M1 141 - 129 oe
(b) A1 12 cao
(b) Answer: 12 cm
(c) B1 the error made the range far larger than it should be (1221 instead of 12), because the range only uses the two extreme values, so a single wrong value distorts it heavily oe
(c) Answer: The error made the range hugely inflated (1221 cm instead of the true 12 cm), because the range depends only on the biggest and smallest values, so one incorrect extreme value badly distorts it.
Question 5
(a) B1 30 cao
(a) Answer: 30 minutes
(b) B1 55 cao
(b) Answer: 55 minutes
(c) B1 55 is much greater than (isolated from) the rest of the data, which is clustered between 28 and 34 minutes oe
(c) Answer: 55 minutes is far higher than every other runner's time, which are all clustered between 28 and 34 minutes.
Question 6
(a) B1 38 cao
(a) Answer: 38 degrees C
(b) B1 a genuine, unusually hot day / short heatwave could have occurred oe
(b) Answer: There may genuinely have been an unusually hot day (e.g. a brief heatwave), so the reading could be correct.
(c) B1 the value may be a genuine, real reading, and removing it without checking would lose important information / make the data set unrepresentative; it should only be removed once checked against other records oe
(c) Answer: 38 degrees C might be a genuine reading, so removing it straight away could hide a real weather event and make the data set less accurate; it should only be removed if checking confirms it was an error.
Question 7
(a) B1 any sensible strategy, e.g. re-weigh puppy C to get the true value, or estimate/impute using the mean of the other puppies' weights oe
(a) Answer: Re-weigh puppy C to find the true value, or estimate the missing weight using the mean of the other 7 puppies.
(b) B1 loses information about puppy C / reduces the sample size, which could make the data set less representative oe
(b) Answer: Deleting the record loses all information about puppy C and reduces the sample size, which could make the results less representative.
(c) M1 (4.2+3.9+4.5+4.0+3.8+4.1+4.3) / 7 oe
(c) A1 4.11 kg (awrt 4.1) cao
(c) Answer: 4.11 kg (awrt 4.1 kg)
Question 8
(a) B1 inconsistent capitalisation, e.g. Blue, BLUE and blue are all recorded differently oe
(a) B1 inconsistent spelling / typing errors, e.g. 'Blu' and 'Greeen' instead of 'Blue' and 'Green' oe
(a) Answer: The responses use inconsistent capitalisation (Blue/BLUE/blue) and inconsistent spelling (Blu, Greeen).
(b) M1 correctly groups all the spelling/capitalisation variants into the three colours oe
(b) A1 Blue = 9, Green = 6, Red = 5, all correct
(b) Answer: Blue = 9, Green = 6, Red = 5
Question 9
(a) M1 5+6+5+7+6+5+6+7+45 (= 92), oe
(a) A1 mean = £10.22 (awrt 10.22) cao
(a) B1 median = £6 cao
(a) Answer: Mean = £10.22 (awrt), median = £6
(b) M1 5+6+5+7+6+5+6+7 (= 47), oe
(b) A1 mean = £5.875 (= £5.88 to 2 dp) cao
(b) B1 median = £6 cao
(b) Answer: Mean = £5.875 (= £5.88 to 2 dp), median = £6
(c) B1 the mean changed a lot (from £10.22 to £5.88 approx) but the median stayed the same (£6), showing the median is much less affected by (more resistant to) an outlier than the mean oe
(c) Answer: The mean fell a lot (from about £10.22 to about £5.88) but the median stayed exactly the same (£6), showing the median is far less affected by an outlier than the mean.
Question 10
(a) B1 process cao
(a) Answer: The process stage
(b) B1 if inconsistent or incorrect data is analysed, the averages/conclusions calculated would be wrong or misleading, so cleaning first ensures the interpretation is based on accurate, comparable data oe
(b) Answer: If Jasmine analyses uncleaned data, her averages and conclusions could be wrong or misleading, so the data must be made accurate and consistent before it is interpreted.
(c) B1 remove/exclude them from the data set, since they contain no usable information oe
(c) Answer: Remove the 3 blank questionnaires from the data set, since they contain no usable information.
(d) B1 guessed values are not genuine data and could bias/distort the results, as they may not reflect what those pupils actually do oe
(d) Answer: Guessed values would not be genuine data and could bias the results, since they may not reflect what those pupils actually do.
Question 11
(a) M1 groups the 10 rows by Member ID oe
(a) A1 5 cao
(a) Answer: 5 unique members
(b) B1 M1042 cao
(b) B1 3 (times) cao
(b) Answer: M1042, recorded 3 times
Question 12
(a) B1 0.4 and 0.3 millilitres would be an impossibly small amount of coffee, so these must have been recorded in litres instead oe
(a) Answer: 0.4 ml and 0.3 ml would be far too small for a cup of coffee, so these values must actually be in litres, not millilitres.
(b) M1 multiplies each value by 1000 oe
(b) A1 400 ml and 300 ml, both correct
(b) Answer: Cup B = 400 ml, Cup D = 300 ml
Question 13
(a) B1 year-month-day (ISO) format oe
(a) Answer: Year-month-day format (2026-07-11)
(b) B1 a computer or reader could misread which part is the day and which is the month, so the appointments could be sorted into the wrong order oe
(b) Answer: A date such as 04/07/2026 could be misread as the 4th of July or as April 7th depending on the format expected, so mixed formats could sort the appointments in the wrong order.
(e) B1 yes, because 15 is greater than the upper boundary of 8 oe
(e) Answer: Yes, 15 is an outlier because it is greater than the upper boundary of 8.
(f) B1 boundaries marked at 0 and 8 (ft), and 15 correctly circled as the outlier
(f) Answer: Boundaries marked at 0 and 8; the value 15 is circled as the outlier.
Question 16
(a) B1 any sensible strategy, e.g. contact volunteer 3 to get the true reading, or estimate it using the mean/median of the other readings oe
(a) Answer: Contact volunteer 3 to get the true reading, or estimate the missing value using the mean of the other readings.
(b) B1 690 bpm is far higher than a resting heart rate could ever realistically be, so a digit has probably been added, or a decimal point missed, by mistake oe
(b) B1 69 bpm cao
(b) Answer: 690 bpm is impossible for a resting heart rate, so it is a data-entry error; the true value was probably 69 bpm.
(c) B1 the volunteer may have just been exercising or feeling anxious, which would genuinely raise their heart rate; 132 bpm, unlike 690 bpm, is a realistic value for a human heart rate oe
(c) Answer: The volunteer may have just been exercising or feeling anxious, which would genuinely raise their heart rate; unlike 690 bpm, 132 bpm is a realistic reading.
(d) B1 690 is impossible and must be corrected/treated as an error, whereas 132 is a possible genuine reading and should be kept (and investigated) rather than deleted or changed oe
(d) Answer: Row 5 (690 bpm) is impossible and must be corrected as an error, but row 8 (132 bpm) is a possible genuine reading, so it should be kept rather than deleted or changed in the same way.
(e) M1 68+71+70+70+69+72+69+132+70 (= 691), oe
(e) A1 76.8 bpm (awrt 76.8) cao
(e) Answer: 76.8 bpm (awrt)
(f) B1 including 132 bpm raises the mean noticeably (from about 69.9 bpm to about 76.8 bpm), showing that even a genuine unusual value can pull the mean up considerably oe
(f) Answer: Including 132 bpm raises the mean from about 69.9 bpm to about 76.8 bpm, showing that even a genuine extreme value can noticeably affect the mean.