Statistics: Data Presentation and Interpretation Depth
In A-level Statistics, Data Presentation Depth extends the core measures of location, spread and correlation to the settings that catch students out under exam pressure: reading a histogram where the class widths are not equal (so height alone does not represent frequency), linearly coding a data set to simplify a mean or standard deviation calculation and then reversing the coding, and describing skewness by comparing the mean and the median. The examiner is testing whether a student knows that a histogram's AREA, not its height, represents frequency, can correctly transform a mean and a standard deviation under a coding, and can justify a description of skew rather than just naming it.
Before you start
Make sure you're comfortable with these topics first:
Method
- Decide what the question needs. A histogram question needs frequency density = frequency / class width, remembering that AREA represents frequency. A coding question needs the given transformation (e.g. y = (x-a)/b) applied to, or reversed from, the mean and standard deviation. A skewness question needs a comparison of the mean and the median (or the quartiles).
- To read a histogram, use frequency = frequency density x class width for each bar. If one class's frequency (or frequency density) is unknown, use a given total frequency and subtract the known classes' frequencies to find it, then divide by that class's width.
- To estimate the median, a quartile, or a frequency for a range that does not align with the class boundaries, use linear interpolation: assume the data is spread evenly across the relevant class, and take the appropriate proportion of that class's width (for a median or quartile) or of its frequency (for an estimated count within part of a class).
- For coded data with y = (x-a)/b: mean(y) = (mean(x) - a) / b, and standard deviation(y) = standard deviation(x) / b (subtracting the constant a does not change the spread, only dividing by b does). Rearrange either way, depending on which mean or standard deviation the question gives you.
- To describe skewness, compare the mean and the median. If mean > median, the data is positively skewed (a longer tail of unusually high values pulls the mean up); if mean < median, the data is negatively skewed (a longer tail of unusually low values pulls the mean down); if mean is approximately equal to median, the data is roughly symmetrical.
- Always finish an interpretation question in the words of the context given (what a difference in spread, or a skew, actually means for these particular data), rather than leaving a bare mathematical statement.
Worked example
A histogram represents the time, in minutes, spent by 90 customers in a shop, with these class widths and frequency densities: 0 to 10 minutes, frequency density 1.2; 10 to 20 minutes, frequency density 3.0; 20 to 40 minutes, frequency density 1.4; 40 to 70 minutes, frequency density d (unknown). Given that 90 customers were surveyed in total, find the value of d, and hence find the number of customers who spent at least 40 minutes in the shop.
- Recall that, in a histogram, frequency = frequency density x class width for each bar.
- Find the frequency of each known class: 0-10 minutes: 1.2 x 10 = 12. 10-20 minutes: 3.0 x 10 = 30. 20-40 minutes: 1.4 x 20 = 28.
- Sum the known frequencies: 12 + 30 + 28 = 70.
- Since the total is 90, the unknown (40-70 minute) class has frequency 90 - 70 = 20.
- This class has width 70 - 40 = 30, so d = 20/30 = 2/3 (about 0.67).
- Final answer: d = 2/3, and 20 customers spent at least 40 minutes in the shop.
Practice questions
Try each question, then tap to reveal the answer.
Q1A histogram bar represents a class of width 5 and is drawn with frequency density 4. State the frequency represented by this bar.Show answer
Answer: Frequency = frequency density x class width = 4 x 5 = 20.
Q2A data set of exam marks, x, is coded using y = (x-40)/2. Given that the mean of y is 12.5, find the mean of x.Show answer
Answer: mean(x) = 2 x mean(y) + 40 = 2(12.5) + 40 = 65.
Q3Using the same coding y = (x-40)/2, given that the standard deviation of y is 3, find the standard deviation of x.Show answer
Answer: standard deviation(x) = 2 x standard deviation(y) = 2(3) = 6.
Q4A data set has mean 45 and median 52. State, with a reason, the type of skew shown.Show answer
Answer: Negative skew, since the mean (45) is lower than the median (52) - a tail of unusually low values is pulling the mean down below the median.
Q5A histogram represents the mass, in kg, of parcels at a depot. The class 10 to 15 kg contains 18 parcels and is drawn with frequency density 3.6. Verify this frequency density is consistent with the class width, and state the height that the bar for a different class, of width 8 kg containing 24 parcels, should be drawn at.Show answer
Answer: 18/5 = 3.6, consistent with the given frequency density. The second bar should be drawn at frequency density 24/8 = 3.
Q6The times, in seconds, taken by 40 students to complete a puzzle are grouped: 0-20 seconds (8 students), 20-30 seconds (17 students), 30-50 seconds (15 students). Use linear interpolation to estimate the median time.Show answer
Answer: The median (20th value) lies in the 20-30 class, since 8 students are below it and 8+17=25 are below the next boundary. Median = 20 + ((20-8)/17) x 10 = 27.1 seconds (3 s.f.).
Q7Explain why, on a histogram with unequal class widths, you cannot compare two classes' frequencies just by comparing the heights of their bars.Show answer
Answer: Frequency is represented by the AREA of a bar (frequency density x class width), not by its height alone. Two classes with different widths need different heights to represent the same frequency, so height alone is only directly comparable when every class has the same width.
Exam-style questions
Written in the style of a A Level Maths exam paper, with a full mark scheme.
A histogram summarises the distance travelled to work, in km, by 120 office workers, with these class widths and frequency densities: 0-5 km, frequency density 4.4; 5-10 km, frequency density 9.2; 10-20 km, frequency density 2.9; 20-40 km, frequency density h. (a) Find the frequency for each of the first three classes. (3) (b) Given that 120 workers were surveyed in total, find the value of h. (2) (c) Using linear interpolation, estimate the median distance travelled, giving your answer to 3 significant figures. (3)
Show mark scheme
Tick each line you got. Your score builds from the marks on the scheme.
Nothing ticked yet - 8 available
The times, in seconds, taken by a group of athletes to run 100 m are coded using y = 10(x-9), giving mean(y) = 8.5 and standard deviation(y) = 4. (a) Find the mean and standard deviation of the original times, x. (3) (b) Given that the median of the original times is 9.05 seconds, determine, with a reason, the type of skewness shown by the data. (2)
Show mark scheme
Tick each line you got. Your score builds from the marks on the scheme.
Nothing ticked yet - 5 available
A histogram represents the age, in years, of 90 members of a gym, with these class widths and frequency densities: 16-25 years, frequency density 2; 25-30 years, frequency density 9; 30-50 years, frequency density 1.35. (a) Show that these frequency densities are consistent with a total of 90 members. (2) (b) Assuming ages are spread uniformly within the 25-30 class, estimate the number of members aged between 27 and 30. (2)
Show mark scheme
Tick each line you got. Your score builds from the marks on the scheme.
Nothing ticked yet - 4 available
Free printable worksheet
Want more practice on paper? Download the statistics: data presentation and interpretation depth worksheet pack - 11 pages of exam-style questions with a full mark scheme. One email opens every download in this browser for 14 days - no account, no card. Print it for personal and classroom use.
Next topics
Not quite what you needed?
Tell us what is missing on statistics: data presentation and interpretation depth, or which topic to write up next. Every request is read, and we reply to every one.
Build a full practice pack.
This topic is one of hundreds in the library - pick the ones a student needs and generate a printable PDF in minutes.