Admissions tests / ESAT / Maths 1 / Statistics and probability

Test standard. 15 questions, 15 marks, about 27 minutes.

ESAT Mathematics 1: Statistics and probability, set 2

Tables, charts and diagrams, averages and spread, sampling, probability of single and combined events, tree and Venn diagrams and conditional probability.

Download the questions (PDF) Download with worked solutions (PDF)

  • Answer all questions. No calculator.
  • Each question has exactly one correct answer.
  1. 11 mark

    A Maths resit was taken by 200 students in two classes.

    Class X: 30 passed and 20 failed (a class total of 50). Class Y: 80 passed and 70 failed (a class total of 150). Overall totals: 110 passed and 90 failed, out of the grand total of 200 students.

    A student is chosen at random from those who PASSED the resit. What is the probability that this student is from Class Y?

    1. A 2/5
    2. B 7/9
    3. C 8/11
    4. D 8/15
  2. 21 mark

    A vertical line chart shows the number of pets owned by each of 15 families.

    The chart has a vertical line of height 1 at 0 pets, height 2 at 1 pet, height 3 at 2 pets, height 6 at 3 pets, height 2 at 4 pets, and height 1 at 5 pets.

    Work out the interquartile range of the number of pets owned.

    1. A 1
    2. B 3
    3. C 5
    4. D 8
  3. 31 mark

    A histogram has two bars showing the ages of people at an event. Bar P covers 10 <= age < 15 (a class width of 5) with frequency density 3.6. Bar Q covers 15 <= age < 25 (a class width of 10) with frequency density 2.1.

    Which bar represents more people, and by how many?

    1. A Bar P represents more people, by 3.
    2. B Bar P represents more people, by 1.5.
    3. C Bar Q represents more people, by 39.
    4. D Bar Q represents more people, by 3.
  4. 41 mark

    A cumulative frequency graph for the length, in minutes, of 80 phone calls passes through the points (0,0), (5,14), (10,34), (15,55), (20,65) and (25,80), where each point is (time, cumulative frequency).

    Use the graph to estimate the interquartile range of the call lengths.

    1. A 15
    2. B 11
    3. C 12.5
    4. D 17.5
  5. 51 mark

    Two machines, A and B, each fill 6 bags of flour. The masses, in kg, are: Machine A: 2.0, 2.1, 1.9, 2.0, 2.2, 1.8. Machine B: 2.5, 1.5, 2.0, 2.3, 1.7, 2.0.

    Which statement correctly compares the two machines, using the mean and the range?

    1. A The machines have the same average mass, but Machine A's masses are more consistent (less spread out).
    2. B Machine B has the higher average mass, and Machine B's masses are more consistent.
    3. C The machines have the same average mass, and Machine B's masses are more consistent.
    4. D Machine A has the higher average mass, but Machine B's masses are more consistent.
  6. 61 mark

    The lengths, in cm, of a sample of pencils are grouped into two class intervals. The interval 0 <= L < 10 has a frequency of 6 pencils. The interval 10 <= L < 20 has a frequency of x pencils.

    The estimated mean length of all the pencils is 12 cm. Work out the value of x.

    1. A 1.5
    2. B 2.8
    3. C 14
    4. D 24
  7. 71 mark

    A scientist plots the number of ice creams sold (y) against the daily maximum temperature (x) for 20 summer days, and finds a strong positive correlation.

    Which conclusion can correctly be drawn from this scatter graph alone?

    1. A The graph proves that hot weather causes more ice cream sales.
    2. B As the temperature increases, the number of ice creams sold tends to decrease.
    3. C Since there is a strong correlation, temperature and ice cream sales must be completely unrelated.
    4. D As the temperature increases, the number of ice creams sold tends to increase, but the graph alone cannot prove that higher temperatures cause the increase in sales.
  8. 81 mark

    A scatter graph plots the cost, in pounds, of catering for a party (y) against the number of guests (x). The line of best fit has equation y = 4x + 60, valid for x-values within the range of the data collected.

    Using this line, estimate how many guests were catered for if the total cost was 180 pounds.

    1. A 30
    2. B 45
    3. C 60
    4. D 116
  9. 91 mark

    A frequency tree records data for 240 customers at a cafe. Of these, 150 order a hot drink and the rest order a cold drink. Of the hot-drink customers, 36 also order a cake. Of the cold-drink customers, 27 also order a cake.

    A customer is chosen at random from those who ordered a HOT drink. What is the probability that this customer also ordered a cake?

    1. A 3/20
    2. B 6/25
    3. C 9/50
    4. D 21/50
  10. 101 mark

    A fair spinner has 8 equal sections, numbered 1 to 8. It is spun 200 times.

    Based on the assumption that the spinner is fair, how many times would you expect it to land on a number that is a multiple of 3?

    1. A 75
    2. B 25
    3. C 0.25
    4. D 50
  11. 111 mark

    A dice is suspected of being biased. It is rolled 30 times and lands on 6 eight times.

    A fair dice would be expected to land on 6 with theoretical probability 1/6. Which statement is best supported by this data?

    1. A Since 8/30 is close to 1/6, this data supports the conclusion that the dice is fair.
    2. B The theoretical probability of rolling a 6 is 8/30, and the relative frequency observed is 1/6.
    3. C The relative frequency (8/30) is higher than the theoretical probability for a fair dice (1/6), which suggests the dice might be biased towards landing on 6, but 30 rolls may not be enough trials to be sure.
    4. D The dice must be biased, because the relative frequency (8/30) is not exactly equal to the theoretical probability (1/6).
  12. 121 mark

    A bag contains balls that are red, blue, yellow or green. A ball is picked at random. P(red) = 0.2, P(blue) = 0.15, P(yellow) = 0.4, P(green) = 0.25.

    Work out the probability that the ball is either red or yellow.

    1. A 0.08
    2. B 0.6
    3. C 1.0
    4. D 0.35
  13. 131 mark

    In a class of 40 students, 22 study French, 17 study Spanish, and 9 study both French and Spanish.

    Using a Venn diagram, work out how many students study NEITHER French nor Spanish.

    1. A 1
    2. B 19
    3. C 13
    4. D 10
  14. 141 mark

    A fair coin is flipped and, separately, a fair four-sided spinner (numbered 1 to 4) is spun.

    By listing all equally likely outcomes in a possibility space, find the probability of getting Heads on the coin AND an even number on the spinner.

    1. A 1/4
    2. B 1/3
    3. C 3/4
    4. D 1/2
  15. 151 mark

    A box contains 4 red pens and 6 blue pens, 10 pens in total. A pen is picked at random, its colour noted, and then it is put back in the box before a second pen is picked at random (i.e. with replacement).

    Work out the probability that both pens picked are red.

    1. A 2/15
    2. B 4/5
    3. C 4/25
    4. D 2/5

Worked solutions

Every question below carries the reasoning, not just the answer. The official material for this test publishes a correct option letter and nothing else.

  1. Question 1Answer: C

    1. The question restricts attention to students who PASSED, so the denominator must be the total number of passes, which is 30 + 80 = 110.
    2. Within those who passed, the number from Class Y is 80 (read directly from the table).
    3. The probability is therefore 80 out of 110, written as the fraction 80/110.
    4. Dividing numerator and denominator by 10 gives 80/110 = 8/11, which does not simplify further since 11 is prime and does not divide 8.
    5. So the answer is 8/11.
    • Why not A: This uses the grand total of all 200 students (giving 80/200 = 2/5) instead of restricting the denominator to only those who passed.
    • Why not B: This reads off the FAIL column instead of the PASS column, using Class Y's 70 fails out of the 90 total fails (70/90 = 7/9), rather than Class Y's passes out of the total passes.
    • Why not D: This uses Class Y's own row total (150) as the denominator instead of the total number who passed (110), effectively finding P(Pass given Class Y) instead of the requested P(Class Y given Pass).
  2. Question 2Answer: A

    1. Listing the 15 pet-counts in order from the given frequencies gives: 0, 1, 1, 2, 2, 2, 3, 3, 3, 3, 3, 3, 4, 4, 5.
    2. With 15 values, the lower quartile is at position (15 + 1)/4 = 4, and the upper quartile is at position 3 x (15 + 1)/4 = 12.
    3. The 4th value in the list is 2, so the lower quartile is 2. The 12th value in the list is 3, so the upper quartile is 3.
    4. The interquartile range is the upper quartile minus the lower quartile: 3 - 2 = 1.
    5. So the interquartile range of the number of pets owned is 1.
    • Why not B: This is the median (the value at the middle position of the ordered list), not the interquartile range; the median here is the 8th value in the list, which is 3.
    • Why not C: This is the range (the highest value minus the lowest value, 5 - 0 = 5), not the interquartile range.
    • Why not D: This subtracts the RANK POSITIONS of the upper and lower quartiles (the 12th position minus the 4th position, 12 - 4 = 8) instead of subtracting the pet-count VALUES found at those positions.
  3. Question 3Answer: D

    1. On a histogram, frequency (the number of people) equals class width multiplied by frequency density, not the frequency density alone.
    2. Bar P covers ages 10 <= age < 15, a width of 5, with frequency density 3.6, so its frequency is 5 x 3.6 = 18.
    3. Bar Q covers ages 15 <= age < 25, a width of 10, with frequency density 2.1, so its frequency is 10 x 2.1 = 21.
    4. Comparing the two frequencies, Bar Q (21) represents more people than Bar P (18).
    5. The difference is 21 - 18 = 3, so Bar Q represents more people, by 3.
    • Why not A: This correctly computes the difference in frequencies (21 - 18 = 3) but names the wrong bar as having more people; Bar Q's frequency (21) is actually larger than Bar P's (18).
    • Why not B: This subtracts the frequency densities directly (3.6 - 2.1 = 1.5) and assumes the bar with the taller frequency density (Bar P) has more people, without multiplying either frequency density by its class width to find the actual frequency.
    • Why not C: This adds the two bars' frequencies together (18 + 21 = 39) instead of subtracting them to find the difference.
  4. Question 4Answer: B

    1. With 80 calls, the lower quartile is at cumulative frequency n/4 = 20, and the upper quartile is at 3n/4 = 60.
    2. Between (5,14) and (10,34), cumulative frequency rises from 14 to 34 (an increase of 20) as time rises from 5 to 10 minutes. The target of 20 is (20-14)/20 = 6/20 = 3/10 of the way through this rise, giving a lower quartile of 5 + (3/10 x 5) = 5 + 1.5 = 6.5 minutes.
    3. Between (15,55) and (20,65), cumulative frequency rises from 55 to 65 (an increase of 10) as time rises from 15 to 20 minutes. The target of 60 is (60-55)/10 = 1/2 of the way through this rise, giving an upper quartile of 15 + (1/2 x 5) = 15 + 2.5 = 17.5 minutes.
    4. The interquartile range is the upper quartile minus the lower quartile: 17.5 - 6.5 = 11.
    5. So the interquartile range is 11 minutes.
    • Why not A: This reads off the class boundaries containing the quartiles (the interval's upper boundary of 20 for the upper quartile, and the lower boundary of 5 for the lower quartile) instead of interpolating within each interval, giving 20 - 5 = 15.
    • Why not C: This finds 25% and 75% of the TIME axis (which runs from 0 to 25 minutes), giving 6.25 and 18.75, instead of finding the times at which the CUMULATIVE FREQUENCY reaches 20 and 60.
    • Why not D: This correctly finds the upper quartile (17.5 minutes) but stops there, forgetting to subtract the lower quartile to get the interquartile range.
  5. Question 5Answer: A

    1. Machine A's masses sum to 2.0+2.1+1.9+2.0+2.2+1.8 = 12.0, so its mean is 12.0/6 = 2.0 kg.
    2. Machine B's masses sum to 2.5+1.5+2.0+2.3+1.7+2.0 = 12.0, so its mean is also 12.0/6 = 2.0 kg. The two machines have the same average mass.
    3. Machine A's range is its highest value minus its lowest value: 2.2 - 1.8 = 0.4 kg.
    4. Machine B's range is 2.5 - 1.5 = 1.0 kg. Since 0.4 is less than 1.0, Machine A's masses are less spread out, i.e. more consistent.
    5. So the correct comparison is: the machines have the same average mass, but Machine A's masses are more consistent.
    • Why not B: This claims Machine B has the higher average, but both machines have the same mean mass of 2.0 kg; it also claims Machine B is more consistent, but Machine B's range (1.0 kg) is larger than Machine A's (0.4 kg), meaning Machine A is the more consistent machine.
    • Why not C: This correctly identifies that the average masses are equal, but reverses the consistency comparison: Machine A's range (0.4 kg) is smaller than Machine B's (1.0 kg), so Machine A is the more consistent machine, not Machine B.
    • Why not D: This claims Machine A has the higher average mass, but the two means are actually equal at 2.0 kg; it also reverses the consistency comparison, since Machine A's smaller range (0.4 kg) makes it the more consistent machine.
  6. Question 6Answer: C

    1. For grouped data, each class interval is represented by its midpoint: the interval 0 <= L < 10 has midpoint 5, and the interval 10 <= L < 20 has midpoint 15.
    2. The estimated mean is (sum of frequency x midpoint) divided by (total frequency). The known class contributes 6 x 5 = 30, and the unknown class contributes x x 15 = 15x, so the total is 30 + 15x, over a total frequency of 6 + x.
    3. Setting the mean equal to 12 gives the equation (30 + 15x)/(6 + x) = 12.
    4. Multiplying both sides by (6 + x) gives 30 + 15x = 12(6 + x) = 72 + 12x.
    5. Rearranging: 15x - 12x = 72 - 30, so 3x = 42, giving x = 14.
    • Why not A: This substitutes each interval's UPPER boundary (10 and 20) for its midpoint (5 and 15) in the mean equation, instead of using the correct midpoints, solving 6 x 10 + 20x = 12(6+x) for x.
    • Why not B: This keeps the total number of pencils fixed at 6 (the known frequency) when solving, forgetting that the total must also include the unknown frequency x, solving 6x5+15x = 12x6 instead of 12x(6+x).
    • Why not D: This leaves out the known interval's own contribution to the weighted total (6 pencils x midpoint 5 = 30) from the left-hand side of the equation, effectively solving 15x = 12(6+x) instead of 30+15x = 12(6+x).
  7. Question 7Answer: D

    1. A positive correlation means that as one variable (temperature) increases, the other variable (ice cream sales) tends to increase too.
    2. The scatter graph shows this tendency, but correlation alone does not establish causation: a third factor, or coincidence, could explain the pattern.
    3. So the graph supports the statement that sales tend to rise with temperature, while explicitly NOT supporting the stronger claim that temperature causes the rise.
    4. This matches the statement that sales tend to increase with temperature, but causation cannot be proved from the graph alone.
    • Why not A: A strong correlation shows that two variables tend to change together, but it does not by itself prove that one causes the other; other factors could be involved, so 'proves' overstates what a scatter graph can show.
    • Why not B: A strong POSITIVE correlation means the two variables tend to increase together, not that one decreases as the other increases; a decrease would describe a negative correlation instead.
    • Why not C: A strong correlation means the variables show a clear, consistent pattern together; it is evidence of a strong relationship, not evidence of 'no relationship' at all.
  8. Question 8Answer: A

    1. The line of best fit is y = 4x + 60, where y is the total cost and x is the number of guests.
    2. Substituting the known cost gives 180 = 4x + 60.
    3. Subtracting the intercept from both sides: 4x = 180 - 60 = 120.
    4. Dividing both sides by the gradient: x = 120/4 = 30.
    5. So the line of best fit estimates that 30 guests were catered for.
    • Why not B: This divides the total cost of 180 by the gradient of 4 directly (180/4 = 45), without first subtracting the fixed cost, the y-intercept of 60.
    • Why not C: This adds the y-intercept to the cost before dividing ((180+60)/4 = 60), instead of subtracting it.
    • Why not D: This correctly subtracts the intercept (180 - 60 = 120) but then subtracts the gradient again instead of dividing by it (120 - 4 = 116).
  9. Question 9Answer: B

    1. The question restricts attention to customers who ordered a HOT drink, so the denominator is the total number of hot-drink customers, which is 150.
    2. Of these 150 hot-drink customers, 36 also ordered a cake, read directly from the hot-drink branch of the tree.
    3. The probability is 36 out of 150, i.e. 36/150.
    4. Dividing numerator and denominator by 6 gives 36/150 = 6/25.
    5. So the probability is 6/25.
    • Why not A: This uses the total number of customers (240) as the denominator (36/240 = 3/20) instead of restricting to only those who ordered a hot drink (150).
    • Why not C: This uses the number of COLD-drink customers who also ordered a cake (27) instead of the number of HOT-drink customers who did (36), reading the wrong branch of the tree.
    • Why not D: This adds together everyone who ordered a cake, from both hot-drink and cold-drink customers (36+27 = 63), instead of using only the hot-drink branch's cake figure (36).
  10. Question 10Answer: D

    1. The spinner has 8 equally likely sections, numbered 1 to 8. The multiples of 3 within this range are 3 and 6, which is 2 out of the 8 sections.
    2. The probability of landing on a multiple of 3 in a single spin is 2/8, which simplifies to 1/4.
    3. The expected number of times this happens in 200 spins is found by multiplying the probability by the number of trials: 200 x 1/4.
    4. 200 x 1/4 = 50.
    5. So you would expect the spinner to land on a multiple of 3 about 50 times.
    • Why not A: This counts 3, 6 and 9 as multiples of 3, but the spinner only has sections numbered 1 to 8, so 9 is not a possible outcome; only 3 and 6 actually qualify.
    • Why not B: This counts only 3 as a multiple of 3 within the range 1 to 8, overlooking that 6 is also a multiple of 3.
    • Why not C: This correctly finds the probability of landing on a multiple of 3 (2/8 = 1/4 = 0.25) but stops there, reporting the probability itself instead of multiplying by 200 to find the expected number of spins.
  11. Question 11Answer: C

    1. For a fair, unbiased dice, the theoretical probability of rolling a 6 is 1/6, since each of the 6 faces is equally likely.
    2. The experiment gives a relative frequency of 8/30 for rolling a 6, which is higher than 1/6.
    3. 8/30 is approximately 0.267 and 1/6 is approximately 0.167, so the difference is noticeable rather than negligible.
    4. This difference suggests the dice might be biased towards 6, but with only 30 rolls the sample is small, so the difference could still be due to chance rather than genuine bias.
    5. A much larger number of rolls would be needed to be more confident about whether the dice is actually biased.
    • Why not A: 8/30 is approximately 0.267, while 1/6 is approximately 0.167; these are not close, so this option misjudges how different the observed relative frequency is from the theoretical probability.
    • Why not B: This swaps the two ideas around: the THEORETICAL probability for a fair dice is 1/6 (based on the dice having 6 equally likely faces), while the RELATIVE FREQUENCY from the experiment is 8/30 (based on what was actually observed).
    • Why not D: A relative frequency from a limited number of trials is expected to vary somewhat from the theoretical probability even for a genuinely fair dice, simply due to chance; a small difference after only 30 rolls does not, by itself, prove bias.
  12. Question 12Answer: B

    1. A ball cannot be more than one colour at once, so 'red' and 'yellow' are mutually exclusive events.
    2. For mutually exclusive events, the probability of either one occurring is found by adding their individual probabilities.
    3. P(red or yellow) = P(red) + P(yellow) = 0.2 + 0.4.
    4. 0.2 + 0.4 = 0.6.
    5. So the probability that the ball is either red or yellow is 0.6.
    • Why not A: This multiplies the two probabilities (0.2 x 0.4 = 0.08) as if 'red or yellow' required both events to combine like independent events joined by AND, instead of adding them, since a ball cannot be both red and yellow at once (mutually exclusive events).
    • Why not C: This adds all four colours' probabilities together (0.2+0.15+0.4+0.25 = 1.0), which only confirms that the probabilities are exhaustive; it does not answer the question about red or yellow specifically.
    • Why not D: This adds the probability of red to the probability of BLUE (0.2+0.15 = 0.35) instead of to the probability of yellow, reading the wrong colour from the question.
  13. Question 13Answer: D

    1. The number of students who study French only is 22-9 = 13 (the 9 who study both are subtracted so they are not double-counted).
    2. The number who study Spanish only is 17-9 = 8.
    3. The number who study at least one of the two subjects is the sum of French only, Spanish only, and both: 13+8+9 = 30.
    4. The number who study neither subject is the class size minus this total: 40-30 = 10.
    5. So 10 students study neither French nor Spanish.
    • Why not A: This treats French and Spanish as if no student studied both, simply subtracting both totals from the class size (40-22-17 = 1), ignoring that the 9 students who study both are being subtracted twice by this method.
    • Why not B: This subtracts the 9 students who study both TWICE from the combined total (22+17-9-9 = 21 studying at least one, so 40-21 = 19), when the overlap should only be subtracted once to avoid double-counting.
    • Why not C: This is the number of students who study French ONLY (22-9 = 13), not the number who study neither subject.
  14. Question 14Answer: A

    1. The coin has 2 equally likely outcomes and the spinner has 4 equally likely outcomes, so the combined possibility space has 2x4 = 8 equally likely outcomes in total.
    2. The even numbers on the spinner are 2 and 4, so the outcomes satisfying 'Heads AND an even number' are (Heads,2) and (Heads,4), which is 2 outcomes.
    3. The probability is 2 out of 8 equally likely outcomes, i.e. 2/8.
    4. 2/8 simplifies to 1/4.
    5. So the probability of Heads and an even number is 1/4.
    • Why not B: This adds the number of coin outcomes and spinner outcomes (2+4 = 6) to find the total sample space, instead of multiplying them (2x4 = 8), since each of the 2 coin outcomes can occur alongside each of the 4 spinner outcomes.
    • Why not C: This finds the probability of 'Heads OR an even number' (which includes all 4 Heads outcomes plus the 2 Tails-and-even outcomes, giving 6 out of 8) instead of 'Heads AND an even number', which requires both conditions together.
    • Why not D: This ignores the coin entirely and just finds the probability of an even number on the spinner alone (2/4 = 1/2), instead of the combined probability of Heads on the coin AND an even number on the spinner.
  15. Question 15Answer: C

    1. Since the first pen is replaced before the second pick, the two picks are independent events, each with the same probability.
    2. For each pick, P(red) = 4/10, which simplifies to 2/5.
    3. Because the picks are independent, the probability that both are red is found by multiplying: P(both red) = 2/5 x 2/5.
    4. 2/5 x 2/5 = 4/25.
    5. So the probability that both pens picked are red is 4/25.
    • Why not A: This treats the second pick as if the first pen were NOT replaced, using 3 red pens out of 9 remaining for the second draw (4/10 x 3/9 = 2/15); but the pen IS put back before the second pick, so the second draw still has 4 red out of 10.
    • Why not B: This adds the two picks' probabilities instead of multiplying them (4/10+4/10 = 4/5), treating 'both red' as if it meant something found by combining the two picks with addition rather than requiring both independent events to occur together.
    • Why not D: This finds only the probability that the FIRST pen is red (4/10 = 2/5) and stops there, without accounting for the second pick at all.

More on this strand

More free ESAT practice

Every strand of the published ESAT specification, with worked solutions throughout.