Data Cleaning and Outliers (Higher)
Data cleaning is the process of checking a data set for errors, such as impossible values, typing mistakes or duplicate entries, and correcting or removing them before analysis. An outlier is a value that lies unusually far from the rest of the data, and part of data cleaning involves deciding whether an outlier is a genuine, if extreme, result or an error that should be corrected or removed.
Before you start
Make sure you're comfortable with these topics first:
Method
- Scan the data set for values that are impossible or clearly wrong for context, such as a negative age or a temperature of 500 degrees C.
- Check for duplicate entries or obvious typing errors, such as a decimal point in the wrong place.
- Identify outliers using a rule, such as any value more than 1.5 x IQR below the lower quartile or above the upper quartile, or a value stated as an outlier by the question.
- Decide whether each outlier is a genuine extreme result, which should be kept but noted, or a data-entry error, which should be corrected if the true value is known or removed if not.
- Recalculate any summary statistics, such as the mean or range, after cleaning, since these are sensitive to errors and outliers.
- State clearly which values were removed or changed and why, since data cleaning decisions should be justified rather than made silently.
Worked example
A nurse records the heights, in cm, of 9 children as: 132, 128, 135, 130, 129, 1310, 133, 131, 128. Identify the likely error in this data set, suggest a correction, and explain your reasoning.
- All the values are between about 128 and 135 cm except for 1310, which is far higher than the rest and not a realistic height for a child.
- 1310 is likely a typing error, for example an extra digit added or a misplaced decimal point (131.0 recorded without the point).
- The most likely correct value is 131, since this fits closely with the surrounding data (128 to 135 cm).
- So the final answer is: 1310 should be corrected to 131 cm, as it is far outside the realistic range for the other children's heights and is very likely a data-entry error.
Practice questions
Try each question, then tap to reveal the answer.
Q1A data set of ages is: 12, 13, 12, 14, -12, 13. Which value is clearly an error, and why?Show answer
Answer: -12, because a negative age is impossible.
Q2A set of weekly pocket money amounts, in pounds, is: 5, 6, 5, 500, 7, 6. Identify the likely outlier.Show answer
Answer: 500 pounds - far higher than the rest, likely a data-entry error (perhaps meant to be 5.00).
Q3Test scores out of 50 are: 32, 28, 35, 30, 29, 31, 95. Which score cannot be correct, and why?Show answer
Answer: 95, because the test is out of 50, so a score of 95 is impossible.
Q4A data set has lower quartile Q1 = 20 and upper quartile Q3 = 32. Using the rule 'outlier if more than 1.5 x IQR beyond Q1 or Q3', find the outlier boundaries.Show answer
Answer: Lower boundary = 2, upper boundary = 50 (IQR = 12, 1.5 x IQR = 18, so 20-18 and 32+18).
Q5Using the boundaries from the previous question (2 and 50), state whether a value of 55 in the data set would be classed as an outlier.Show answer
Answer: Yes, 55 is above the upper boundary of 50, so it is an outlier.
Q6A data set of 10 house prices, in thousands of pounds, has a mean of 210 without an outlier of 900 included, but a mean of 279 when the 900 is included. Explain why a statistician might exclude the 900 when reporting a typical house price.Show answer
Answer: The mean is very sensitive to extreme values; including 900 pulls the mean well above where most prices actually lie, so excluding it (or using the median) gives a more typical figure.
Exam-style questions
Written in the style of a GCSE Statistics exam paper, with a full mark scheme.
The masses, in kg, of 8 new-born lambs are recorded as: 4.2, 4.5, 3.9, 4.1, 4.3, 41, 4.0, 4.4. (a) Identify the value that is likely to be a data-entry error. (b) Suggest a corrected value, giving a reason.
Show mark scheme
Tick each line you got. Your score builds from the marks on the scheme.
Nothing ticked yet - 3 available
A data set of 11 delivery times, in minutes, has lower quartile Q1 = 18 and upper quartile Q3 = 30. (a) Calculate the interquartile range (IQR). (b) A statistician defines an outlier as any value more than 1.5 x IQR above Q3 or below Q1. Calculate the two boundary values used to identify outliers. (c) A delivery time of 55 minutes is recorded. State whether this would be classed as an outlier.
Show mark scheme
Tick each line you got. Your score builds from the marks on the scheme.
Nothing ticked yet - 4 available
The following are the annual salaries, in thousands of pounds, of 7 employees at a small company: 28, 31, 29, 30, 32, 250, 27. (a) Explain why the value 250 should be investigated before being included in any analysis. (b) Calculate the mean salary with the value 250 included. (c) Calculate the mean salary with the value 250 excluded. (d) Comment on which mean better represents a typical employee's salary.
Show mark scheme
Tick each line you got. Your score builds from the marks on the scheme.
Nothing ticked yet - 5 available
See real GCSE Statistics past-paper questions, with official mark schemes →
Free printable worksheet
Want more practice on paper? Download the data cleaning and outliers (higher) worksheet pack - 18 pages of exam-style questions with a full mark scheme. One email opens every download in this browser for 14 days - no account, no card. Print it for personal and classroom use.
This topic is chapter 11 of GCSE Statistics Higher Workbook 1, the whole course as one free printable PDF.
Next topics
Not quite what you needed?
Tell us what is missing on data cleaning and outliers (higher), or which topic to write up next. Every request is read, and we reply to every one.
Build a full practice pack.
This topic is one of hundreds in the library - pick the ones a student needs and generate a printable PDF in minutes.