Statistics: Data Presentation and Interpretation Depth - Worksheets, Questions and Revision

12 original exam-style questions - 6 pages of questions with a full mark scheme - free printable PDF.

Download PDFJump to mark scheme (page 7)
« Previous: Statistics: Hypothesis TestingNext: Statistics: Probability Depth »
Revision Library
revisionlibrary.co.uk
A-Level · Statistics

S6 Statistics: Data Presentation and Interpretation Depth

EDEXCEL 9MA0 · Calculator allowed · about 135 minutes
Total Marks
Name: _______________________________    Date: ____ / ____ / ______
Answer ALL questions. Show all your working.

A-Level Statistics: Data Presentation and Interpretation - Depth Practice

Original material written for Revision Library, styled on Edexcel AS/A-level Mathematics (9MA0) Statistics content.

This pack extends core data presentation and interpretation skills to full A-level depth: measures of location and spread (including coding), identifying outliers by the quartile rule and the standard deviation rule, box plots and skewness, histograms and frequency density, linear interpolation, correlation and regression (including coding and residuals), and evaluating statistical models in context, in the style of the large data set used across the qualification. All contexts and figures are original.

1
A tutor records the daily rainfall, in mm, at a weather station near Leeds over 8 consecutive days in March:

12, 15, 9, 21, 18, 14, 11, 16
(a)Calculate the mean and the standard deviation of the rainfall for these 8 days, giving each answer to 3 significant figures where appropriate.(4)
(b)The 8 rainfall values are converted from millimetres to centimetres. State, without further calculation, the new mean and the new standard deviation.(2)
(Total for Question 1 is 6 marks)
2
The times, in minutes, taken by 15 students to complete an S6 statistics quiz are recorded and arranged in ascending order:

8, 10, 11, 12, 12, 13, 14, 14, 15, 16, 16, 17, 18, 19, 31
(a)Find the lower quartile, the median and the upper quartile of this data set.(3)
(b)A value is defined as an outlier if it lies more than 1.5 x IQR beyond the nearer quartile. Using your answers to part (a), determine whether the value 31 is an outlier, showing your working clearly.(3)
(Total for Question 2 is 6 marks)
3
The times, in minutes, taken by students to complete the same homework task were summarised for two schools using the following five-figure summaries.

School A: minimum = 5, Q1 = 18, median = 25, Q3 = 30, maximum = 34
School B: minimum = 9, Q1 = 15, median = 17, Q3 = 26, maximum = 40
(a)Calculate the interquartile range for each school.(2)
(b)By comparing the position of the median relative to the quartiles (or the lengths of the whiskers) for each school, describe and justify the skewness of each distribution.(4)
(c)Using an appropriate measure, compare the overall variation (spread) in completion times between the two schools, and comment on whether your two measures give the same conclusion.(2)
(Total for Question 3 is 8 marks)
4
The masses, in kg, of a sample of parcels processed at a sorting depot are summarised in the grouped frequency table below. One frequency is missing.

Mass, m (kg): 0 < m ≤ 10, frequency 8
Mass, m (kg): 10 < m ≤ 20, frequency 22
Mass, m (kg): 20 < m ≤ 30, frequency f (unknown)
Mass, m (kg): 30 < m ≤ 50, frequency 18
Mass, m (kg): 50 < m ≤ 80, frequency 12

It is known that the frequency density for the class 20 < m ≤ 30 is 2.6.
(a)Find the value of f.(2)
(b)Calculate the frequency density for each of the other four classes, and state, with a reason, which class would be represented by the tallest bar if a histogram were drawn for this data.(3)
(c)Estimate the number of parcels with mass greater than 30 kg.(1)
(Total for Question 4 is 6 marks)
5
The table shows the distribution of the number of hours, h, spent revising per week by 60 sixth-form students during study leave.

Hours, h: 0 ≤ h < 5, midpoint 2.5, frequency 6
Hours, h: 5 ≤ h < 10, midpoint 7.5, frequency 14
Hours, h: 10 ≤ h < 15, midpoint 12.5, frequency 20
Hours, h: 15 ≤ h < 20, midpoint 17.5, frequency 12
Hours, h: 20 ≤ h < 25, midpoint 22.5, frequency 8

A coding y = (x - 12.5)/5 is used, where x is the class midpoint.
(a)Complete the coded midpoint values, y, for each class.(2)
(b)Using the coded data, calculate an estimate for the mean number of hours spent revising per week.(3)
(c)Calculate an estimate for the standard deviation of the number of hours spent revising per week.(4)
(Total for Question 5 is 9 marks)
6
For a sample of 12 UK secondary schools, x represents the average class size and y represents the mean GCSE point score per pupil. The following summary statistics were calculated:

n = 12, Sxx = 248, Syy = 196, Sxy = -152
(a)Calculate the product moment correlation coefficient, r, between x and y.(3)
(b)Interpret the value of r found in part (a) in context.(2)
(c)The headteacher of one school claims that 'smaller class sizes cause higher GCSE point scores'. Comment on the validity of this claim, referring to the value of r found in part (a).(2)
(Total for Question 6 is 7 marks)
7
A researcher models the relationship between the number of hours of sunshine, h, and the maximum daily temperature, t degrees Celsius, recorded at a UK coastal town over a 15-day period in July. The regression line of t on h is found to be

t = 14.2 + 0.85h
(a)Interpret the value 0.85 in the context of the model.(2)
(b)Interpret the value 14.2 in the context of the model, and comment on the reliability of this interpretation.(2)
(c)Use the regression equation to estimate the maximum temperature on a day with 6 hours of sunshine.(2)
(d)Explain why it would not be appropriate to use this equation to estimate the maximum temperature on a day with 20 hours of sunshine.(2)
(e)One value of h in the data set was recorded as 0.2 hours, with an actual temperature of 20.8 degrees Celsius. Calculate the residual for this data point, and comment on what this suggests.(3)
(Total for Question 7 is 11 marks)
8
The number of texts sent in a day by 9 students is recorded:

22, 25, 19, 30, 27, 21, 26, 23, 95
(a)Calculate the mean and standard deviation of all 9 values.(4)
(b)The value 95 is identified as a recording error and removed. Calculate the new mean and standard deviation for the remaining 8 values.(4)
(c)Comment on the effect that removing this value has had on the mean and on the standard deviation.(2)
(Total for Question 8 is 10 marks)
9
The daily maximum temperatures, in degrees Celsius, were recorded for a sample of 31 days in May at two UK weather stations, Aviemore and Camborne, in the style of the large data set. Summary statistics are given below.

Aviemore (n = 31): minimum = 6.1, Q1 = 9.4, median = 11.8, Q3 = 14.2, maximum = 19.5
Camborne (n = 31): minimum = 10.8, Q1 = 12.6, median = 13.9, Q3 = 15.7, maximum = 18.0
(a)Calculate the interquartile range for each location.(2)
(b)Using the rule that an outlier is a value more than 1.5 x IQR beyond the nearer quartile, calculate the upper outlier boundary for Aviemore, and state whether the maximum value of 19.5 degrees Celsius is an outlier.(3)
(c)Compare and contrast the distributions of maximum daily temperature at the two locations, referring to the median and the interquartile range.(3)
(d)Explain one reason, in context, why Aviemore might show a greater spread of daily maximum temperatures than Camborne.(1)
(Total for Question 9 is 9 marks)
10
The table below shows the distribution of the ages, in complete years, of 200 members of a gym. Ages are recorded to the nearest whole year, so the class 16-25 has boundaries 15.5 and 25.5, and so on.

Age (years): 16-25, frequency 30
Age (years): 26-35, frequency 54
Age (years): 36-45, frequency f (unknown)
Age (years): 46-55, frequency 38
Age (years): 56-70, frequency 20
(a)Show that the frequency, f, for the class 36-45 is 58.(1)
(b)Calculate the frequency density for each class (using class widths based on the true class boundaries), and state, with a reason, which class would be represented by the tallest bar if a histogram were drawn for this data.(2)
(c)Use linear interpolation to estimate the median age of the gym members.(3)
(d)Estimate the number of gym members aged under 30.(3)
(Total for Question 10 is 9 marks)
11
A student, Kwame, is investigating the relationship between the number of GCSE subjects studied at higher tier, x, and the total revision hours per week, y, for a sample of 10 students. To simplify the calculations, the data is coded using u = x - 8 and v = y - 20, giving the following summary statistics for the coded data:

n = 10, sum(u) = 14, sum(v) = 35, sum(u2) = 54, sum(v2) = 205, sum(uv) = 97
(a)Show that Suv, the value of Sxy for the coded data, is equal to 48.(2)
(b)Given that Suu = 34.4, find the equation of the regression line of v on u in the form v = a + bu, giving a and b to 3 significant figures.(4)
(c)Hence find the equation of the regression line of y on x, in the form y = c + dx, giving c and d to 3 significant figures.(3)
(d)Use your equation from part (c) to estimate the total revision hours per week for a student studying 10 GCSE subjects at higher tier.(2)
(Total for Question 11 is 11 marks)
12
As part of a study styled on the large data set, a researcher records the daily mean windspeed, in knots, at a UK weather station for the month of October (31 days). The summary statistics are sum(x) = 217 and sum(x2) = 1859.4 (n = 31). The five smallest values in the data set are 3, 4, 4, 5, 5 knots, and the data also include one exceptionally high reading of 42 knots, recorded during a named storm.
(a)Calculate the mean and standard deviation of the windspeed data.(4)
(b)Using the rule that a value is an outlier if it lies more than 2 standard deviations from the mean, determine whether the value of 42 knots should be classified as an outlier.(3)
(c)The researcher is deciding whether to remove the value of 42 knots before calculating summary statistics intended to describe a 'typical' October day. Discuss, with reference to the context and to your answers in parts (a) and (b), whether this value should be removed from the data set.(6)
(Total for Question 12 is 13 marks)
Mark scheme · S6 Statistics: Data Presentation and Interpretation Depth

Question 1

Question 2

Question 3

Question 4

Question 5

Question 6

Question 7

Question 8

Question 9

Question 10

Question 11

Question 12