Statistics: Data Presentation and Interpretation - Worksheets, Questions and Revision

13 original exam-style questions - 5 pages of questions with a full mark scheme - free printable PDF.

Download PDFJump to mark scheme (page 6)Read the revision guide
« Previous: Statistics: SamplingNext: Statistics: Probability »
Revision Library
revisionlibrary.co.uk
A-Level · Statistics

S2 Statistics: Data Presentation and Interpretation

EDEXCEL 9MA0 · Calculator allowed · about 130 minutes
Total Marks
Name: _______________________________    Date: ____ / ____ / ______
Answer ALL questions. Show all your working.
1
A company has 850 employees split across three departments: Admin (170 employees), Production (510 employees) and Sales (170 employees). The company wants to survey employees about a proposed new flexible working policy.
(a)State what is meant, in the context of sampling, by a 'sampling frame'.(1)
(b)The company decides to take a stratified sample of 40 employees, proportional to department size. Calculate the number of employees that should be sampled from each department.(3)
(c)Describe how systematic sampling could instead be used to select 40 employees from a numbered list of all 850 employees, and state one advantage this method has over simple random sampling in this context.(3)
(Total for Question 1 is 7 marks)
2
A cafe owner records, for 8 different days last winter, the outdoor temperature and the number of hot chocolates sold.
(a)State the type of correlation you would expect between outdoor temperature and the number of hot chocolates sold.(1)
(b)Describe the general shape of the scatter diagram you would expect to see for this data, referring to the direction of the trend.(1)
(c)The owner also wants to investigate the relationship between outdoor temperature and the number of cold drinks sold. State the type of correlation expected and explain your reasoning.(2)
(d)On one unusually cold day in December, a local event caused hot chocolate sales to be much higher than the temperature that day would otherwise suggest. Explain the effect this data point is likely to have on the product moment correlation coefficient (PMCC) calculated for temperature against hot chocolate sales across the whole data set.(2)
(Total for Question 2 is 6 marks)
3
The times, in minutes, taken by 9 runners to complete a 5 km training route are: 34, 36, 29, 41, 38, 33, 30, 45, 37.
(a)Calculate the mean time.(2)
(b)Calculate the standard deviation of these times.(3)
(c)A tenth runner joins the group with a time of 60 minutes, a clear outlier. Without recalculating, state and justify the effect that including this runner's time would have on the mean and on the standard deviation.(2)
(Total for Question 3 is 7 marks)
4
The typing speeds, in words per minute, of 11 trainees on a data entry course are: 42, 45, 47, 48, 50, 51, 53, 55, 58, 60, 90.
(a)Calculate the mean and the standard deviation of these 11 typing speeds.(4)
(b)Show that the value 90 is an outlier, using the rule that a value is an outlier if it lies more than 2 standard deviations from the mean.(3)
(c)Given that, without the outlier, the sum of the remaining 10 typing speeds is 509 and the sum of their squares is 26201, calculate the new mean and new standard deviation for the remaining 10 trainees.(4)
(d)Comment on the effect that removing the outlier has had on the mean and on the standard deviation of the typing speeds.(2)
(Total for Question 4 is 13 marks)
5
In a mathematics exam, marks x are coded using y = (x - 30)/5. For 15 students, the coded marks y satisfy: sum of y = 24, sum of y2 = 140.
(a)Calculate the mean of the coded marks, y.(1)
(b)Hence calculate the mean exam mark, x.(2)
(c)Calculate the standard deviation of the coded marks, y, and hence calculate the standard deviation of the original exam marks, x.(4)
(Total for Question 5 is 7 marks)
6
The heights, in cm, of 50 sunflower plants grown in a trial were recorded in the grouped frequency table below.
Height (cm)150-160160-170170-180180-190190-200
Frequency5122094
(a)Using the midpoint of each class, calculate an estimate of the mean height.(3)
(b)Using linear interpolation, estimate the median height.(3)
(c)Using your answer to part (a), calculate an estimate of the variance of the heights.(3)
(Total for Question 6 is 9 marks)
7
The time, t minutes, taken by 100 children to complete a jigsaw puzzle was recorded in the grouped frequency table below.
Time t (minutes)0≤t<1010≤t<2020≤t<3535≤t<5050≤t<70
Frequency1228f1510
(a)Given that a total of 100 children completed the puzzle, calculate the value of f.(2)
(b)On a histogram of this data, calculate the frequency density that should be used for the bar representing 20≤t<35.(2)
(c)Using linear interpolation, estimate the median time taken to complete the puzzle.(3)
(d)Assuming the times are evenly spread within the class 50≤t<70, estimate the number of children who took more than 60 minutes to complete the puzzle.(3)
(Total for Question 7 is 10 marks)
8
The weekly hours worked by 13 part-time staff at a supermarket, listed in ascending order, are: 12, 15, 15, 18, 20, 21, 22, 24, 25, 27, 29, 32, 55.
(a)Find the median, the lower quartile (Q1) and the upper quartile (Q3) of these weekly hours.(3)
(b)Show that 55 hours is an outlier, using the rule that a value is an outlier if it lies more than 1.5 x IQR beyond the nearer quartile.(3)
(c)State the five figures needed to draw a box plot for this data, and explain how the outlier at 55 should be represented on the diagram.(3)
(d)Using your values for the median, Q1 and Q3, comment on the skewness of the distribution of weekly hours.(2)
(Total for Question 8 is 11 marks)
9
Two classes, A and B, sat the same maths test, out of 50 marks. Box plots of their results gave the following five-figure summaries.
Class A: minimum = 10, Q1 = 18, median = 26, Q3 = 40, maximum = 48
Class B: minimum = 15, Q1 = 27, median = 30, Q3 = 33, maximum = 36
(a)Compare the medians and the interquartile ranges of the two classes, in the context of the test.(3)
(b)Using the quartiles, comment on the skewness of each class's distribution of marks.(2)
(c)A teacher claims 'Class B had less variation in test marks than Class A.' By considering both the range and the interquartile range, comment on whether this claim is fair.(3)
(Total for Question 9 is 8 marks)
10
For 8 days last spring, a shopkeeper recorded the number of hours of sunshine, x, and the shop's ice cream sales, y, in hundreds of pounds.
Hours of sunshine, x: 2, 3, 5, 6, 7, 8, 9, 10
Ice cream sales, y (GBP hundreds): 1.2, 1.8, 2.6, 3.0, 3.4, 4.0, 4.5, 4.9
You may use: sum(x) = 50, sum(y) = 25.4, sum(xy) = 184.1, sum(x2) = 368, sum(y2) = 92.26
(a)Calculate the value of the product moment correlation coefficient (PMCC), r, for these data.(3)
(b)Interpret the value of r found in part (a), in context.(1)
(c)Find the equation of the regression line of y on x in the form y = a + bx, giving a and b to 3 significant figures.(4)
(d)Use your regression line to estimate the ice cream sales on a day with 6.5 hours of sunshine, and comment on the reliability of this estimate.(3)
(Total for Question 10 is 11 marks)
11
House prices, in thousands of pounds, were recorded for two streets. For 20 houses in Elm Street: sum(x) = 4200, sum(x2) = 920000. For 15 houses in Birch Street: sum(x) = 3300, sum(x2) = 746000.
(a)Calculate the mean and the standard deviation of house prices in Elm Street.(4)
(b)Calculate the mean and the standard deviation of house prices in Birch Street.(4)
(c)Using your answers to parts (a) and (b), compare house prices in Elm Street and Birch Street.(2)
(Total for Question 11 is 10 marks)
12
A distribution of delivery times has mean = 45, median = 50, mode = 58 and standard deviation = 12. It is also known that Q1 = 38, Q2 = 50 and Q3 = 54.
(a)Calculate Pearson's coefficient of skewness, given by (mean - mode)/standard deviation.(2)
(b)Calculate the quartile coefficient of skewness, given by (Q3 + Q1 - 2Q2)/(Q3 - Q1).(2)
(c)Interpret the values found in parts (a) and (b) in terms of the skewness of the distribution.(2)
(d)State one advantage of using the quartile coefficient of skewness, found in part (b), rather than Pearson's coefficient of skewness, found in part (a).(1)
(Total for Question 12 is 7 marks)
13
A researcher recorded, for 10 UK towns, the number of ice cream shops, x, and the number of sunburn cases reported at local hospitals over one summer, y. The PMCC between x and y was found to be r = 0.91, based on values of x ranging from 2 to 15.
(a)The researcher concludes: 'Opening more ice cream shops causes more cases of sunburn.' Comment on this conclusion, suggesting a variable that could explain the high correlation observed.(3)
(b)State one condition that should be checked before using the regression line from this data to predict y for a town with x = 25 ice cream shops, and explain why this matters here.(2)
(c)A large data set recorded daily mean temperature and daily rainfall for a UK location across a whole year. Explain why it would not be appropriate to calculate a single PMCC using all of this data if you wanted to investigate the relationship between temperature and rainfall during summer months only.(2)
(Total for Question 13 is 7 marks)
Mark scheme · S2 Statistics: Data Presentation and Interpretation

Question 1

Question 2

Question 3

Question 4

Question 5

Question 6

Question 7

Question 8

Question 9

Question 10

Question 11

Question 12

Question 13