Further Statistics: Hypothesis Testing, Type I and Type II Errors - Worksheets, Questions and Revision

12 original exam-style questions - 5 pages of questions with a full mark scheme - free printable PDF.

Download PDFJump to mark scheme (page 6)
« Previous: Further Statistics: Chi-Squared Tests (Goodness of Fit and Contingency Tables)Next: Further Statistics: Discrete Distributions Depth »
Revision Library
revisionlibrary.co.uk
A-Level · Further Statistics 1 (Hypothesis Testing for a Poisson or Binomial Parameter)

FP.FS4 Further Statistics: Hypothesis Testing, Type I and Type II Errors

AQA 7367 · Calculator allowed · about 120 minutes
Total Marks
Name: _______________________________    Date: ____ / ____ / ______
Answer ALL questions. Show all your working.

Hypothesis Testing for a Poisson or Binomial Parameter

Revision Library

For a hypothesis test on a Poisson mean lambda or a binomial probability p, a null hypothesis H0 states a fixed value of the parameter, and an alternative hypothesis H1 states that the parameter has increased, has decreased (a one-tailed test), or is simply different (a two-tailed test). A test statistic X, assumed to follow the distribution stated in H0, is compared against a critical region: the set of values of X for which H0 is rejected in favour of H1.

Because X is discrete, the probability of X falling in a proposed critical region can rarely be made to equal a chosen significance level exactly. The usual convention is to choose the largest critical region whose probability does not exceed the significance level requested. This probability is called the actual significance level of the test, and it is also equal to P(Type I error), the probability of wrongly rejecting a true H0.

A Type II error occurs when H0 is not rejected even though H1 is actually true. Its probability depends on the true value of the parameter, and is found by calculating the probability that X falls outside the critical region, using the true (alternative) value of the parameter in place of the value stated in H0. The power of a test is defined as 1 - P(Type II error): the probability of correctly rejecting a false H0. Power increases as the true parameter value moves further from the value stated in H0, and is also affected by the significance level chosen and the sample size used.

1
A car insurance company models the number of claims made per month by a certain class of driver using a Poisson distribution, and regularly carries out hypothesis tests on the mean claim rate.
(a)State what is meant by the critical region (rejection region) of a hypothesis test.(1)
(b)Explain why, for a test statistic with a discrete distribution such as Poisson or binomial, the actual significance level of a test is not usually exactly equal to the significance level originally requested.(2)
(c)A test is carried out, at a chosen significance level, of H0: the mean monthly claim rate is unchanged, against H1: the mean monthly claim rate has increased. Define, in the context of this test, a Type I error and a Type II error.(2)
(d)State an equation linking the power of a hypothesis test to P(Type II error).(1)
(Total for Question 1 is 6 marks)
2
A call centre in Preston models the number of customer complaints received per day, under normal conditions, by X ~ Po(4). Following a change to the automated phone system, the manager wants to test, at the 5% significance level, whether the mean number of complaints per day has decreased.
(a)State suitable null and alternative hypotheses, then find the critical region for the test and state the actual significance level.(3)
(b)On a particular day following the change, the number of complaints received is 0. State, with a reason, the conclusion of the test.(2)
(c)Suppose that the change has in fact reduced the true mean number of complaints per day to 2. Find the probability that this test results in a Type II error.(3)
(d)Hence state the power of the test in this case.(1)
(Total for Question 2 is 9 marks)
3
A garden centre in Exeter finds that, historically, 30% of tomato seedlings supplied by a particular grower are affected by blight in damp weather. A random sample of 20 seedlings is taken, and X is the number affected. A new fungicide is applied to the grower's stock, and the centre wants to test, at the 5% significance level, whether the fungicide has reduced the proportion affected below 0.3.
(a)State suitable hypotheses, then find the critical region for the test and state the actual significance level.(4)
(b)After treatment, 1 of the 20 sampled seedlings is found to be affected by blight. State, with a reason, the conclusion of the test.(2)
(c)Suppose that the fungicide has in fact reduced the true proportion of seedlings affected to 0.15. Find the probability that the test results in a Type II error.(3)
(d)Hence state the power of the test in this case.(1)
(e)State, without further calculation, what would happen to the probability of a Type II error found in part (c) if the significance level of the test were increased from 5% to 10%. Justify your answer.(2)
(Total for Question 3 is 12 marks)
4
The number of minor faults found on newly manufactured circuit boards from a particular production line is modelled, under normal conditions, by X ~ Po(6). Quality control engineers want to test, at the 10% significance level, whether the mean number of faults per board has changed from 6 (a two-tailed test).
(a)State suitable hypotheses, then find the critical region for this two-tailed test and state the actual significance level.(6)
(b)A randomly chosen board is found to have 10 faults. State, with a reason, the conclusion of the test.(2)
(c)Suppose that a fault in the manufacturing process has in fact increased the true mean number of faults per board to 9. Find the probability that the test results in a Type II error.(3)
(d)Hence state the power of the test in this case.(1)
(e)Comment on the suitability of this test for detecting the increase in mean described in part (c).(1)
(Total for Question 4 is 13 marks)
5
The number of typing errors per page of a large legal document is modelled, under normal conditions, by X ~ Po(10). A proofreading test uses the critical region X ≤ 3 or X ≥ 18 to test whether the mean number of errors per page differs from 10.
(a)Find the probability of a Type I error for this test.(5)
(Total for Question 5 is 5 marks)
6
A machine at a plastics factory in Telford produces components, 20% of which have historically been classed as premium grade. After the machine is recalibrated, an engineer wants to test, at the 5% significance level, whether the proportion of premium-grade components has increased. A random sample of 25 components is taken, and X is the number classed as premium grade.
(a)State suitable hypotheses, then find the critical region for the test and state the actual significance level.(4)
(b)In the sample, 10 components are classed as premium grade. State, with a reason, the conclusion of the test.(2)
(c)Suppose that the recalibration has in fact increased the true proportion of premium-grade components to 0.35. Find the probability that the test results in a Type II error.(4)
(d)Hence state the power of the test in this case.(1)
(e)State, without further calculation, what would happen to the power of the test found in part (d) if the sample size were increased from 25 to 50, with the significance level kept at 5%. Justify your answer.(2)
(Total for Question 6 is 13 marks)
7
A fairground game at a summer fete in Durham is designed so that a player wins on any go with probability 0.4, independently of other goes. The organiser wants to test, at the 10% significance level, whether the true win probability differs from 0.4, based on X, the number of wins in 15 goes.
(a)State suitable hypotheses, then find the critical region for this two-tailed test and state the actual significance level.(6)
(b)In one trial of 15 goes, there are 3 wins. State, with a reason, the conclusion of the test.(2)
(c)State one advantage of using a 10% significance level rather than a 5% significance level for this test, together with a corresponding disadvantage.(1)
(Total for Question 7 is 9 marks)
8
Referring again to the garden centre's fungicide trial in Question 3, where X ~ B(20, 0.3) under H0, suppose instead that the garden centre wants the actual significance level of the one-tailed (decrease) test to not exceed 1%.
(a)Find the resulting critical region, and state the actual significance level.(3)
(b)Explain briefly why this stricter test is less likely to detect a genuine reduction in the proportion of affected seedlings than the 5% test used in Question 3(a).(2)
(Total for Question 8 is 5 marks)
9
A hospital microbiology lab models the number of bacterial colonies per agar plate, under safe conditions, as a Poisson random variable with a known mean. Two proposed testing procedures are available to check whether the mean number of colonies per plate has risen above the safe threshold: Test A uses a 1% significance level, and Test B uses a 10% significance level. Failing to detect a genuine rise in bacterial colonies (a Type II error) could pose a public health risk, while wrongly signalling a rise when none has occurred (a Type I error) would trigger a costly, unnecessary shutdown of the affected ward for deep cleaning.

Discuss, with reference to Type I and Type II errors, which of the two significance levels the lab should adopt, and justify a recommendation.
(Total for Question 9 is 6 marks)
10
A textile mill in Bradford inspects rolls of fabric for flaws. Under normal production, the number of flaws found on a 20m roll is modelled by X ~ Po(8). Following a change of yarn supplier, the mill wants to test, at the 5% significance level, whether the mean number of flaws per roll has decreased.
(a)State suitable hypotheses, then find the critical region for the test and state the actual significance level.(4)
(b)Suppose that, in fact, the new yarn has reduced the true mean number of flaws per roll to 5. Find the probability of a Type II error.(3)
(c)Hence state the power of the test in this case.(1)
(d)The mill's quality manager wants the power of the test to be at least 0.5 when the true mean is 5, without inspecting more than one roll. Suggest one change to the test that could achieve this, and state one disadvantage of your suggestion.(2)
(Total for Question 10 is 10 marks)
11
A market-research company uses a 12-sided spinner which should show each of the numbers 1 to 12 with equal probability. To check whether the spinner is biased towards even numbers, an assistant spins it 12 times and records X, the number of times an even number occurs. If the spinner is fair, X ~ B(12, 0.5). In one trial, an even number occurs on 9 of the 12 spins.
(a)Calculate the two-tailed p-value associated with this result.(4)
(b)State the conclusion of the test at the 5% significance level, with a reason.(1)
(Total for Question 11 is 5 marks)
12
Referring again to the textile mill's test in Question 10, where X ~ Po(8) under H0 and the critical region for the 5% one-tailed (decrease) test was found to be X ≤ 3, suppose instead that the new yarn has reduced the true mean number of flaws per roll to 3 (rather than 5, as in Question 10).
(a)Using the same critical region as in Question 10, find the probability of a Type II error, and hence the power of the test, in this case.(5)
(b)Compare the power found in part (a) with the power found in Question 10(c), and explain, with reference to the two Poisson distributions involved, why the powers differ.(3)
(Total for Question 12 is 8 marks)
Mark scheme · FP.FS4 Further Statistics: Hypothesis Testing, Type I and Type II Errors

Question 1

Question 2

Question 3

Question 4

Question 5

Question 6

Question 7

Question 8

Question 9

Question 10

Question 11

Question 12