GCSE Statistics · Topic guide

Data Cleaning and Outliers (Foundation)

Data cleaning is the process of checking a data set for errors, such as impossible values, typing mistakes or missing entries, and correcting or removing them before analysis. An outlier is a value that is much higher or lower than the rest of the data and does not fit the general pattern. In GCSE Statistics, spotting and dealing with outliers correctly is an important step before calculating averages or drawing graphs.

Foundation tierCollecting DataEdexcelAQA

Before you start

Make sure you're comfortable with these topics first:

Method

  1. Look through the data set for values that are impossible or clearly wrong for the context, such as a negative age or a height of 1000 cm.
  2. Check for likely typing errors, such as a missing decimal point or an extra zero.
  3. Identify outliers by comparing each value to the general spread of the rest of the data.
  4. Decide whether an outlier is a genuine but unusual result or a clear error; only remove values that are errors or that you have good reason to exclude.
  5. Recalculate any averages or statistics using the cleaned data set once errors have been corrected or removed.
  6. State clearly which values were removed or corrected and why, since this affects how the results should be interpreted.

Worked example

A nurse records the weights (kg) of 8 patients: 62, 65, 71, 68, 640, 70, 66, 69. (a) Identify the outlier. (b) Suggest a likely explanation. (c) Calculate the mean weight with the outlier removed.

  1. Compare each value with the rest: 640 kg is far larger than all the other weights (62 to 71 kg), so 640 is the outlier.
  2. A likely explanation is a data entry error, for example the decimal point was missed and the true value was 64.0 kg.
  3. Remove the outlier, leaving 7 values: 62, 65, 71, 68, 70, 66, 69.
  4. Sum the remaining values: 62 + 65 + 71 + 68 + 70 + 66 + 69 = 471.
  5. Divide by the number of values: 471 divided by 7 = 67.285714...
  6. Final answer: the outlier is 640 kg, likely a data entry error, and the mean weight of the remaining 7 patients is 67.3 kg (1 dp).

Practice questions

Type your answer and press Check to be marked straight away, or reveal the answer and mark yourself.

Q1A data set of ages is: 12, 13, 12, 14, 13, 99. Which value is most likely an outlier?Show answer

Answer: 99 (far higher than the rest, likely a recording error for an age).

Got it right?
Q2A shop records daily sales in pounds: 340, 355, 360, 348, 2, 352. Which value looks like a data entry error?Show answer

Answer: 2 (a digit was likely missed, for example it should read 320).

Got it right?
Q3Give one reason you might choose to keep an outlier in a data set rather than remove it.Show answer

Answer: It could be a genuine, correct value that shows real variation, so removing it would lose important information.

Got it right?
Q4A data set of test scores out of 100 is: 55, 60, 58, 62, 59, -10. Which value must be an error, and why?Show answer

Answer: -10, because a test score cannot be negative.

Got it right?
Q5The heights (cm) of 6 plants are: 15, 18, 16, 14, 17, 150. Find the mean of the remaining values after removing the outlier.Show answer

Answer: 16 cm (15 + 18 + 16 + 14 + 17 = 80, 80 divided by 5 = 16).

Got it right?
Q6Explain why the median might be a more appropriate average to use on a data set with an outlier, instead of removing the outlier first.Show answer

Answer: The median is not affected much by extreme values, so it can be used on the original data set without needing to identify or remove an outlier.

Got it right?

Exam-style questions

Written in the style of a GCSE Statistics exam paper, with a full mark scheme.

Q1[3 marks]

The number of pets owned by 7 students is: 1, 2, 1, 0, 2, 1, 15. (a) Identify the outlier in this data set. (b) Explain why it is unlikely to be a genuine value. (c) State one action the researcher could take.

Show mark scheme

Tick each line you got. Your score builds from the marks on the scheme.

Nothing ticked yet - 3 available

Got it right?
Q2[2 marks]

A data set of temperatures in degrees C recorded over a week is: 18, 19, 17, 20, 18, -180, 19. State the outlier and explain briefly why it must be an error rather than a real UK summer temperature.

Show mark scheme

Tick each line you got. Your score builds from the marks on the scheme.

Nothing ticked yet - 2 available

Got it right?
Q3[5 marks]

A researcher records the salaries, in thousands of pounds, of 9 employees: 24, 26, 25, 27, 23, 26, 24, 25, 95. (a) Identify the outlier. (b) Calculate the mean of the original 9 values. (c) Calculate the mean with the outlier removed. (d) Comment on which mean better represents a typical employee's salary.

Show mark scheme

Tick each line you got. Your score builds from the marks on the scheme.

Nothing ticked yet - 5 available

Got it right?

See real GCSE Statistics past-paper questions, with official mark schemes

Free printable worksheet

Want more practice on paper? Download the data cleaning and outliers (foundation) worksheet pack - 16 pages of exam-style questions with a full mark scheme. One email opens every download in this browser for 14 days - no account, no card. Print it for personal and classroom use.

This topic is chapter 1 of GCSE Statistics Foundation Workbook 1, the whole course as one free printable PDF.

Other cuts of this worksheet:

Next topics

Ready to practise data cleaning and outliers (foundation)? Add it to a printable topic pack for this student in the Pack Builder.

Add to my pack

Not quite what you needed?

Tell us what is missing on data cleaning and outliers (foundation), or which topic to write up next. Every request is read, and we reply to every one.

Build a full practice pack.

This topic is one of hundreds in the library - pick the ones a student needs and generate a printable PDF in minutes.