You're viewing free preview questions. Upgrade to access more MYP4 questions.Upgrade
Statistics and Probability

Statistics and Probability — Free MYP4 Mathematics (Extended) Practice Questions

1QuestionApplications in Real-World Risk AssessmentConcept Practice
2 marks~3 minCriterion D
An insurance company estimates that the probability of a house fire in a given region is 5%5\%, based on historical data across thousands of properties. They apply this figure to assess the fire risk for an individual house valued at 200000200\,000 dollars.

Explain one limitation of using this regional probability to assess the fire risk for a specific house, and identify one individual factor that could cause the actual risk to differ from 5%5\%. [2]
Question diagram

Solutions

2QuestionFairness Bias and RandomnessConcept Practice
4 marks~6 minCriterion A
A quality-control technician suspects a six-sided die used in a board-game factory may be biased. The die is rolled 60 times with the following results:

Outcome123456
Frequency810971313
a
Calculate the relative frequency of rolling a 4. [1]
b
Calculate the absolute difference between the relative frequency of rolling a 4 and its theoretical probability. [2]
c
The factory rejects any die for which the absolute difference between the relative frequency and theoretical probability of any outcome exceeds 0.08 after 60 trials. Advise the technician whether this die should be rejected, justifying your answer with a numerical comparison. [1]
Question diagram

Solutions

3QuestionUsing Probability Trees for Multi-Stage EventsConcept Practice
2 marks~3 minCriterion A
A hospital screening programme tests patients in two sequential stages. In Stage 1, the probability that a patient tests positive is 0.40.4. In Stage 2, regardless of the Stage 1 result, the probability that a patient tests positive is 0.70.7.

A patient is selected at random.
a
Construct a probability tree diagram for the two stages, labelling all branches with their probabilities. [1]
b
Calculate P(both stages positive)P(\text{both stages positive}) and interpret what this value means for the hospital's expected follow-up caseload. [1]
Question diagram

Solutions

4QuestionCalculating Mean Median and ModeConcept Practice
2 marks~3 minCriterion A
A school counsellor records the number of siblings for each student in an 11-person advisory group to understand family structures when planning support sessions.

Number of siblings: 0, 0, 1, 1, 1, 2, 2, 2, 3, 4, 5

The counsellor considers a median of 2 or more siblings as an indicator that students may benefit from a shared-experience peer support group.
a
Calculate the median number of siblings. [1]
b
Advise the counsellor whether the peer support group should be formed, justifying your answer using the median. [1]

Solutions

5QuestionRange and Interquartile RangeConcept Practice
4 marks~6 minCriterion A
A school uses the following box-and-whisker plot to report class test scores.

Min=45,Q1=60,Median=75,Q3=85,Max=100\text{Min} = 45, \quad Q_1 = 60, \quad \text{Median} = 75, \quad Q_3 = 85, \quad \text{Max} = 100

The school's performance policy states that a class is considered consistent if its IQR is less than half its range.
a
Calculate the range and the interquartile range (IQR) of the test scores. [2]
b
Explain how the IQR and the range each describe the spread of the test scores differently. [1]
c
Justify whether this class meets the school's consistency condition, using your results from part (a). [1]
Question diagram

Solutions

6QuestionConstructing and Interpreting Box PlotsConcept Practice
2 marks~3 minCriterion A
A school uses box plots to compare class performance on a standardised mathematics test.

The five-number summary for Class 10B is:
Minimum =40= 40, Q1=55Q_1 = 55, Median =70= 70, Q3=85Q_3 = 85, Maximum =100= 100.

The school's policy states that a class shows sufficient spread in the middle 50% of scores only if the IQR exceeds 25 marks.
a
Calculate the interquartile range (IQR) for Class 10B. [1]
b
Justify whether Class 10B meets the school's policy for sufficient spread. [1]
Question diagram

Solutions

7QuestionSet Notation and Venn Diargram NotationsConcept Practice
2 marks~3 minCriterion A
A school records whether each student in a group of 80 studies French (FF) and whether they study Spanish (SS). The results are displayed in the Venn diagram below.

Identify the region of the Venn diagram that represents FSF \cap S, and explain what membership of this region means for a student in the group. [2]
Question diagram

Solutions

8QuestionBox plots and quartilesConcept Practice
5 marks~8 minCriterion A
The box plot below shows the distribution of scores (out of 100) for 40 students on a mathematics test.

The five-number summary is: minimum =45= 45, Q1=58Q_1 = 58, median =71= 71, Q3=82Q_3 = 82, maximum =96= 96.
a
Write down the interquartile range (IQR) of the scores. [1]
b
A student is described as a "low performer" if their score falls below Q1Q_1. Find the number of students in the class who are low performers. [2]
c
The teacher claims: "The majority of students scored above 71." Justify whether this claim is correct, using the definition of the median. [2]
Question diagram

Solutions

9QuestionComparing Experimental and Expected ResultsAssessment Practice
2 marks~3 minCriterion C
A quality-control analyst tests a six-sided die by rolling it 12 times, recording: 2, 5, 1, 3, 6, 2, 4, 1, 3, 5, 2, 6.
a
Deduce the experimental frequency for each outcome and the expected frequency for each outcome if the die is fair.

Outcome123456
Experimental frequency__________________
Expected frequency__________________ [1]
b
The analyst concludes the die is fair because no single outcome was recorded more than 3 times. Critique this conclusion using the experimental results. [1]

Solutions

10QuestionApplications in Real-World Risk AssessmentAssessment Practice
4 marks~6 minCriterion D
An insurance company records the following data for two groups of drivers over one year.

Age groupUnder 25 — claims: 120total drivers: 800
Age groupOver 25 — claims: 80total drivers: 1200


The relative risk of making a claim is defined as

Relative Risk=P(claimunder 25)P(claimover 25)\text{Relative Risk} = \frac{P(\text{claim} \mid \text{under 25})}{P(\text{claim} \mid \text{over 25})}
a
Calculate P(claimunder 25)P(\text{claim} \mid \text{under 25}) and P(claimover 25)P(\text{claim} \mid \text{over 25}). [2]
b
Calculate the relative risk of making a claim for drivers under 25 compared to drivers over 25. [1]
c
The insurance company charges drivers under 25 a premium 2.5 times higher than drivers over 25. Justify whether this premium multiplier is supported by the claim data. [1]
Question diagram

Solutions

11QuestionOrganizing Data for Effective AnalysisAssessment Practice
6 marks~9 minCriterion C
A school records the travel times (in minutes) of 50 students commuting to school. The stem-and-leaf plot below displays the full dataset.

StemLeaves
02 5 7 8 9
11 3 4 6 7 9 9
20 2 3 5 6 8
31 2 4 5 7 8
40 1 3


Key: 1 | 9 represents 19 minutes

The same data are also organised into a grouped frequency table with intervals 0–10, 11–20, 21–30, 31–40, 41–50.
a
Using the stem-and-leaf plot, identify the modal travel time and calculate the range of travel times. [2]
b
Discuss one advantage and one disadvantage of using the grouped frequency table, compared with the stem-and-leaf plot, for determining the modal travel time and analysing the spread of the data. [2]
c
The school is deciding whether to introduce a subsidised bus service for students whose typical travel time falls within the most common 15-minute band. Analyse how changing the grouped frequency table intervals to 0–15, 16–30, 31–45, 46–60 affects the identified modal class, and advise the school whether the interval choice should influence its decision about which students qualify for the bus service. [2]
Question diagram

Solutions

12QuestionSimple discrete data and classificationAssessment Practice
5 marks~8 minCriterion A
A quadratic function is given by f(x)=x26x+kf(x) = x^2 - 6x + k, where kk is a constant.
a
Write down the axis of symmetry of f(x)f(x). [1]
b
Find the value of kk such that f(x)=0f(x) = 0 has exactly one solution. [2]
c
Justify whether a value of k=12k = 12 produces two distinct real solutions to f(x)=0f(x) = 0. [2]
Question diagram

Solutions

13QuestionDesigning Surveys and QuestionnairesAssessment Practice
6 marks~9 minCriterion D
A school board surveys all 1200 students about a proposed new uniform policy. Paper surveys are distributed during morning registration; 144 are returned one week later. Of these, 132 responses are from Grade 11 students. Among all returned surveys, 98 students support the new policy. The board concludes the policy is popular and plans to implement it.
a
Calculate the response rate of this survey. [1]
b
Identify two distinct types of bias present in the survey results, supporting each with a calculation or specific evidence from the data. [3]
c
Advise the board whether the survey data are sufficient to justify implementing the new uniform policy. [2]
Question diagram

Solutions

14QuestionSimple discrete data and classificationAssessment Practice
10 marks~15 minCriterion B
A pattern of dots is arranged in triangular frames. The first three figures are shown in the diagram.

The number of dots DD in Figure nn follows a quadratic rule of the form
D=an2+bn+c.D = an^2 + bn + c.

Figure number nn123
Number of dots DD61528
a
Find the values of aa, bb, and cc. [4]
b
Show that D=2n2+3n+1D = 2n^2 + 3n + 1 can be written in the factored form D=(2n+1)(n+1)D = (2n+1)(n+1), and verify that this form gives the correct number of dots for n=1,2,3n = 1, 2, 3. [3]
c
Justify why the rule must contain an n2n^2 term by referring to the structure of the dot pattern. [3]
Question diagram

Solutions

15QuestionUsing Probability Trees for Multi-Stage EventsAssessment Practice
7 marks~11 minCriterion B
A quality-control engineer tests microchips by running them through a sequence of independent diagnostic checks. At each check, a chip either passes (P) or fails (F), each with equal probability.
a
Write down the number of distinct outcomes after 1 check, after 2 checks, and after 3 checks. Identify the pattern in your sequence. [2]
b
Deduce a formula for the total number of distinct outcomes after nn checks, each with mm equally likely results. Justify your formula using the structure of a probability tree. [2]
c
The engineer runs 4 checks on each chip. Construct a probability tree for this 4-stage process and use your formula to verify the total number of distinct paths. A chip is accepted only if it passes all 4 checks. Advise the engineer whether this acceptance criterion is suitable for a production line, justifying your answer using probability. [3]
Question diagram

Solutions

16QuestionRecognizing and Modeling Dependent EventsAssessment Practice
6 marks~9 minCriterion D
A factory produces electronic components using two sequential machines. Machine 1 produces a defect with probability 0.10.1. If Machine 1 produces a defect, Machine 2 produces a defect with probability 0.80.8; if Machine 1 does not produce a defect, Machine 2 produces a defect with probability 0.050.05. A component is classified as defective if at least one machine produces a defect.
a
Calculate the probability that a randomly chosen component is defective. [2]
b
Calculate the probability that Machine 1 produced a defect, given that a component is defective. [2]
c
The factory's quality manager uses your answer from part (b) to argue that, when a defective component is found, Machine 1 is the more likely source of the defect and should be inspected first. Advise the quality manager whether this inspection strategy is reliable, identifying one assumption of the model that could undermine it in practice. [2]
Question diagram

Solutions

17QuestionReal-Life Applications Games Genetics and Decision-MakingAssessment Practice
4 marks~6 minCriterion C
A genetics laboratory crosses two pea plants and records the phenotypes of 80 offspring.

Tall with green seeds: 30
Tall with yellow seeds: 20
Short with green seeds: 10
Short with yellow seeds: 20
a
Calculate the conditional probability that a randomly selected offspring is tall, given that it has green seeds. [2]
b
A second laboratory claims that height and seed colour are independent traits in this cross. Justify whether the data support this claim. [2]
Question diagram

Solutions

18QuestionCalculating Mean Median and ModeAssessment Practice
8 marks~12 minCriterion B
A city transport authority records daily passenger counts (in thousands) at five stations over one week. The analyst notices that some stations show symmetric distributions while others do not.

Five data sets are provided:

Set A357911
Set B246810
Set C13579
Set D45678
Set E02468
a
Calculate the mean and median for each data set. Record your results as labelled rows, then describe one pattern you observe across all five sets. [3]
b
Justify why, for any symmetric data set with an odd number of terms, the mean and median must be equal, referring to the role of the central term. [3]
c
The analyst considers using the mean to represent a sixth station whose daily counts are 1, 2, 3, 4, 20 (thousands). Calculate the mean and median for this data set, then advise the analyst whether the mean or the median should be reported to the public, justifying your recommendation in context. [2]

Solutions

19QuestionCalculating Mean Median and ModeAssessment Practice
6 marks~9 minCriterion D
A small business employs 10 people with the following monthly salaries (in dollars):

2000, 2200, 2300, 2500, 2800, 3000, 3500, 4000, 5000, 250002000,\ 2200,\ 2300,\ 2500,\ 2800,\ 3000,\ 3500,\ 4000,\ 5000,\ 25000

The highest salary belongs to the CEO.
a
Calculate the mean monthly salary. [2]
b
Calculate the median monthly salary. [2]
c
The business owner claims the company offers "above-average salaries" to attract new employees, using the mean as evidence. Advise a job applicant which measure of central tendency they should rely on when assessing typical pay at this company, and justify your recommendation. [2]
Question diagram

Solutions

20QuestionComparing Data Sets Using AveragesAssessment Practice
8 marks~12 minCriterion C
A small business owner records daily sales (in dollars) over two weeks:

Week 1120135110128142115130
Week 220095105118122108112


The owner claims that Week 1 was better because its mean daily sales are higher.
a
Calculate the mean daily sales for each week. Show your working. [2]
b
Explain why the mean may be a misleading measure for comparing these two weeks. Identify the specific value that distorts the comparison and support your answer with a calculation. [2]
c
Justify the use of the median as a more appropriate measure of central tendency for this comparison. Calculate the median for each week and explain what the results reveal. [2]
d
The owner is considering whether the exceptional sales day in Week 2 could repeat due to a planned promotion. Advise the owner on whether the median alone is sufficient to inform this decision, justifying which measure — or combination of measures — would lead to a better-informed conclusion. [2]
Question diagram

Solutions

21QuestionComparing Variability Between Data SetsAssessment Practice
6 marks~9 minCriterion C
A store manager models daily sales as 120±15120 \pm 15 dollars, predicting that sales will range across a spread of 3030 dollars.

The actual daily sales (in dollars) over 10 days are recorded below.

Day: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10

Sales: 122, 135, 112, 128, 118, 108, 130, 125, 132, 115
a
Calculate the range of the actual sales data. [2]
b
Calculate the interquartile range (IQR) of the actual sales data. [2]
c
Critique the manager's claim that the model is reliable because the predicted spread of 3030 dollars is close to the actual range. Use both the range and IQR in your response. [2]

Solutions

22QuestionBox and Whisker Plots for Spread of DataAssessment Practice
6 marks~9 minCriterion D
A city planner is comparing housing prices in two neighbourhoods, Eastside and Westside, to decide where to invest in affordable housing programmes.

Prices (thousands of dollars):

Eastside: 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280

Westside: 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320
a
Calculate the five-number summary (minimum, Q1Q_1, median, Q3Q_3, maximum) for each neighbourhood. [2]
b
Calculate the interquartile range (IQR) and range for each neighbourhood. [2]
c
Construct box-and-whisker plots for both neighbourhoods on the same scale. [1]
d
The city planner states: "Eastside is the better target for affordable housing investment because its prices are more consistent." Advise the city planner whether the measures of spread support this conclusion, and identify one limitation of relying solely on IQR and range for this decision. [1]
Question diagram

Solutions

23QuestionRange and standard deviationAssessment Practice
6 marks~9 minCriterion B
The diagram shows Figures 1 to 4 of a pattern made from dots arranged in a triangular grid.

Figure 1 has 1 dot, Figure 2 has 3 dots, Figure 3 has 6 dots, and Figure 4 has 10 dots.
a
Write down the number of dots added when going from Figure 3 to Figure 4. [1]
b
Find a rule for the number of dots DD in Figure nn, giving your answer in the form
D=n(n+1)2.D = \frac{n(n+1)}{2}.
Verify your rule for Figure 3. [3]
c
A student claims that Figure 16 is the first figure in the pattern to contain more than 100 dots. Justify whether the student is correct. [2]
Question diagram

Solutions

24QuestionConstructing and Interpreting Box PlotsAssessment Practice
6 marks~9 minCriterion B
A sports analyst records the recovery times (in minutes) of 23 athletes after a training session. The ordered values are:

4,8,12,15,18,21,24,27,30,33,36,39,42,45,48,51,54,57,60,63,66,69,724, 8, 12, 15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, 63, 66, 69, 72

Quartile positions for datasets of size n=7,11,15,19n = 7, 11, 15, 19 are given below.

n=7n = 7: Q1Q_1 at position 2, median at position 4, Q3Q_3 at position 6

n=11n = 11: Q1Q_1 at position 3, median at position 6, Q3Q_3 at position 9

n=15n = 15: Q1Q_1 at position 4, median at position 8, Q3Q_3 at position 12

n=19n = 19: Q1Q_1 at position 5, median at position 10, Q3Q_3 at position 15
a
Deduce a general rule, in terms of nn, for the positions of the median, Q1Q_1, and Q3Q_3 in an ordered dataset of nn values. [2]
b
Using your rules from part (a), construct the five-number summary for the 23 recovery times, showing all position calculations. [2]
c
The analyst claims that the middle 50% of athletes recover within a 36-minute window. Justify whether the interquartile range supports this claim. [2]
Question diagram

Solutions

25QuestionConstructing and Interpreting Box PlotsAssessment Practice
4 marks~6 minCriterion C
A school records the scores of 20 students on a mathematics test. The five-number summary is:

Minimum =32= 32, Q1=55Q_1 = 55, Median =67= 67, Q3=78Q_3 = 78, Maximum =95= 95
a
Construct a box plot for this data on a number line from 30 to 100, clearly marking all five values. [2]
b
The school's pass threshold is the lower quartile. Justify whether the score distribution provides sufficient evidence that the majority of students passed. [2]
Question diagram

Solutions

26QuestionCumulative Frequency Graphs and CurvesAssessment Practice
6 marks~9 minCriterion D
A cumulative frequency curve shows the mathematics exam scores of 200 students. Key points on the curve are given below.

Score405570
Cumulative frequency5590165


A normal distribution model with mean 55 and standard deviation 15 predicts the following cumulative frequencies.

Score405570
Predicted cumulative frequency32100168
a
Calculate the percentage difference between the actual and predicted cumulative frequencies at a score of 40, using the formula

Percentage difference=actualpredictedactual×100%\text{Percentage difference} = \frac{|\text{actual} - \text{predicted}|}{\text{actual}} \times 100\%

Give your answer to one decimal place. [2]
b
The percentage differences at scores 55 and 70 are 11.1% and 1.8% respectively. Analyse how the accuracy of the normal distribution model varies across the lower, middle, and upper ends of the score distribution. [2]
c
The school's assessment policy states that a model is considered reliable only if it produces a percentage difference below 20% at every key score. Advise the school whether the normal distribution model should be used to make decisions about student performance in this exam. [2]
Question diagram

Solutions

27QuestionSolving Problems using tree diagrams and Venn DiagramsAssessment Practice
4 marks~6 minCriterion C
A bag contains 4 red marbles and 6 blue marbles. Two marbles are drawn without replacement. The tree diagram shows the following probabilities:

First draw: P(Red)=410P(\text{Red}) = \dfrac{4}{10}, P(Blue)=610P(\text{Blue}) = \dfrac{6}{10}

Second draw (after Red): P(RedRed)=39P(\text{Red} \mid \text{Red}) = \dfrac{3}{9}, P(BlueRed)=69P(\text{Blue} \mid \text{Red}) = \dfrac{6}{9}

Second draw (after Blue): P(RedBlue)=49P(\text{Red} \mid \text{Blue}) = \dfrac{4}{9}, P(BlueBlue)=59P(\text{Blue} \mid \text{Blue}) = \dfrac{5}{9}
a
State the probability of drawing a red marble second, given that the first marble drawn was red. [1]
b
Show that the two draws are dependent events by comparing the conditional probabilities of drawing a red marble second with the unconditional probability of drawing a red marble first. [2]
c
A quality-control inspector claims that, for a fair sampling process, the colour drawn first should have no influence on subsequent draws. Advise the inspector whether this bag of marbles satisfies that claim when sampling without replacement. [1]
Question diagram

Solutions

28QuestionQualitative handling of probabilityAssessment Practice
5 marks~8 minCriterion B
The diagram shows the first four figures of a pattern made from small equilateral triangles.

Figure 11 shaded triangle0 unshaded triangles.
Figure 21 shaded triangle3 unshaded triangles.
Figure 31 shaded triangle8 unshaded triangles.
Figure 41 shaded triangle15 unshaded triangles.


The total number of small triangles in Figure nn is denoted T(n)T(n).
a
Write down the values of T(1)T(1), T(2)T(2), T(3)T(3), and T(4)T(4). [1]
b
Find a rule for T(n)T(n) in the form T(n)=an2+bT(n) = an^2 + b, where aa and bb are integers to be determined. [2]
c
Justify why the rule T(n)=n2T(n) = n^2 must contain an n2n^2 term by referring to the geometric structure of Figure nn. [2]
Question diagram

Solutions

29QuestionCalculating Combined ProbabilitiesAssessment Practice
4 marks~6 minCriterion D
A meteorologist models rainfall using a Markov chain. The probability of rain on Day 1 is 0.40.4. Given rain today, the probability of rain tomorrow is 0.70.7; given no rain today, the probability of rain tomorrow is 0.20.2.
a
Construct a tree diagram showing all possible weather outcomes for three consecutive days, labelling every branch with its probability. [1]
b
Calculate the probability that it rains on exactly two of the three days. [2]
c
The meteorologist claims this model is reliable enough to issue a public rain warning whenever the probability of rain on Day 3 exceeds 0.50.5. After two consecutive rainy days, advise the meteorologist whether the model alone is sufficient justification for issuing a public warning. [1]
Question diagram

Solutions

30QuestionStem-and-leaf plots and pictogramsAssessment Practice
8 marks~12 minCriterion A
The back-to-back stem-and-leaf plot below shows the scores (out of 100) for two classes on the same mathematics test.

Class A (leaves) \quad Stem \quad Class B (leaves)

9 8 7 5 342 69\ 8\ 7\ 5\ 3 \quad|\quad 4 \quad|\quad 2\ 6
8 6 5 4 2 151 3 5 7 88\ 6\ 5\ 4\ 2\ 1 \quad|\quad 5 \quad|\quad 1\ 3\ 5\ 7\ 8
7 4 2 060 2 4 6 97\ 4\ 2\ 0 \quad|\quad 6 \quad|\quad 0\ 2\ 4\ 6\ 9
3 172 5 83\ 1 \quad|\quad 7 \quad|\quad 2\ 5\ 8
81 4\quad|\quad 8 \quad|\quad 1\ 4

Leaves are read outward from the stem. For example, Class A's row for stem 4 represents the scores 43, 45, 47, 48, 49.
a
Write down the median and the interquartile range (IQR) for Class A. [2]
b
Find the mean and standard deviation for Class B. Give your answers correct to 3 significant figures. [3]
c
A teacher must decide which class to enter into an inter-school mathematics competition. The competition selects students who score above 65. Using your answers from parts (a) and (b), and the distributions visible in the stem-and-leaf plot, evaluate which class is better suited for the competition, and justify your recommendation. [3]
Question diagram

Solutions

31QuestionStem-and-leaf plots and pictogramsAssessment Practice
8 marks~12 minCriterion D
A school library records the number of books borrowed each day over two weeks. The back-to-back stem-and-leaf plot below shows daily borrowing figures for weekdays (Monday–Friday) on the left and weekend days (Saturday–Sunday) on the right. The library was closed on one Monday due to a public holiday; that day is not included, giving 9 weekday values and 4 weekend values.

Weekdays (read right to left)StemWeekends (read left to right)
9 812 5
6 4 320 7
7 5 331
842


Key: 9 | 1 means 19 books (weekdays); 1 | 2 means 12 books (weekends)

The sorted weekday values (9 values) are: 18, 19, 23, 24, 26, 35, 37, 38, 48.
a
State the median number of books borrowed per weekday. [1]
b
The library manager uses the mean of the weekend data as a model for a "typical weekend day." Calculate the mean number of books borrowed per weekend day, giving your answer to the nearest whole number. [2]
c
The missing Monday is now included. Library records show that Mondays consistently have among the highest daily borrowing figures for weekdays. A student claims: "Including the missing Monday will increase the weekday median."

Justify whether the student's claim is correct, showing clearly how the median changes when a value greater than 38 is added. [3]
d
The library manager intends to present this two-week dataset to the school board as evidence of typical daily borrowing throughout the entire school year, in order to justify a budget increase. Evaluate whether the dataset provides sufficient evidence for this purpose. [2]
Question diagram

Solutions

32QuestionData visualizations and infographicsAssessment Practice
8 marks~12 minCriterion C
The table below shows the global number of smartphone users (in millions) for the years 2018 to 2022, together with the values predicted by the linear model U=330t660200,U = 330t - 660\,200, where tt is the year.

Year20182019202020212022
Actual (millions)28003200360039004100
Predicted (millions)27403070340037304060
a
Calculate the percentage error for the year 2022, using the formula percentage error=predictedactualactual×100%.\text{percentage error} = \frac{|\text{predicted} - \text{actual}|}{\text{actual}} \times 100\%. Give your answer correct to 2 significant figures. [2]
b
A student claims the linear model is a good fit because the predicted values are always within 5% of the actual values. Determine whether this claim holds for every year in the table, showing your working for each year. [3]
c
The diagram shows the actual data points and the linear model plotted on the same axes. Assess whether the linear model or a quadratic model would be more appropriate for predicting smartphone users beyond 2022, referring to the trend visible in the residuals and to one limitation of extrapolation. [3]
Question diagram

Solutions

33QuestionStem-and-leaf plots and pictogramsAssessment Practice
10 marks~15 minCriterion B
The diagram shows a growing pattern of dots arranged in triangular frames. Figure 1 has 8 dots, Figure 2 has 15 dots, and Figure 3 has 24 dots.
a
Find the number of dots in Figure 4 and Figure 5. [2]
b
A student claims the number of dots DD in Figure nn can be written in the form
D=an2+bn+cD = an^2 + bn + c
where aa, bb, and cc are integers. Use any three figures to form a system of equations and solve to find the values of aa, bb, and cc. [4]
c
The student says: "Since the second differences of the sequence are constant, the rule must be quadratic, and the coefficient aa tells us how the pattern grows layer by layer." Verify that the second differences are constant and equal to 2, then justify whether the student's claim that the rule must be quadratic is correct. [4]
Question diagram

Solutions

34QuestionLines of best fitAssessment Practice
7 marks~11 minCriterion B
The diagram shows Figures 1, 2 and 3 of a pattern made from unit squares arranged in rectangles.

Figure (nn)123
Squares (SS)3815
a
Write down the number of unit squares in Figure 4. [1]
b
Find a rule for the number of unit squares SS in Figure nn. Show that your rule can be written in the form
S=n(n+2).S = n(n + 2). [3]
c
A student claims that SS is always an odd number when nn is odd. Justify this claim using your rule, then determine whether Figure 13 is the only figure for which S=195S = 195, giving a reason for your answer. [3]
Question diagram

Solutions

35QuestionLines of best fitAssessment Practice
7 marks~11 minCriterion D
A survey records the number of hours studied per week (xx) and the test score (yy, out of 100) for 10 students. A line of best fit is drawn with equation
y=8x+30.y = 8x + 30.
The table below shows predicted and actual scores for three students.

xx (hours)258
Predicted yy467094
Actual yy506588
a
Calculate the percentage error for the student who studied for 5 hours, using
percentage error=predictedactualactual×100%.\text{percentage error} = \frac{|\text{predicted} - \text{actual}|}{\text{actual}} \times 100\%.
Give your answer to 2 significant figures. [2]
b
A student claims she studied for 7 hours and scored 90. Deduce whether her score is above or below the score predicted by the model, and find the exact difference. [2]
c
Another student scored 74 on the test. Using the model, find the number of hours this student is predicted to have studied. Hence advise whether a teacher should use this model to predict the score of a student who studies for 12 hours, justifying your answer with reference to the data given. [3]
Question diagram

Solutions

36QuestionLines of best fitAssessment Practice
6 marks~9 minCriterion C
The scatter diagram below shows test scores (out of 100) and revision times (in hours) for 10 students. A line of best fit has been drawn with equation
y=8x+28y = 8x + 28
where xx is revision time in hours and yy is the predicted test score.

The table below shows actual and predicted scores for three students.

xx (hours)258
Predicted yy446892
Actual yy506580
a
Write down the yy-intercept of the line of best fit and interpret its meaning in this context. [2]
b
Calculate the percentage error for the student who revised for 8 hours, using
percentage error=predictedactualactual×100%.\text{percentage error} = \frac{|\text{predicted} - \text{actual}|}{\text{actual}} \times 100\%.
Give your answer correct to 1 decimal place. [2]
c
A student claims the model is equally reliable for predicting scores at 2 hours and at 8 hours of revision. Using the percentage errors for both x=2x = 2 and x=8x = 8, assess whether this claim is correct. [2]
Question diagram

Solutions

37QuestionLines of best fitAssessment Practice
12 marks~18 minCriterion A
The table below shows the number of hours studied (xx) and the test score (yy) for 8 students.

xx (hours): 2, 3, 5, 6, 8, 9, 11, 14

yy (score): 55, 60, 72, 78, 85, 82, 95, 68
a
State the coordinates of the point that appears to be an outlier, and give one reason why it is unusual. [2]
b
Using all 8 data points, calculate the equation of the line of best fit in the form y=mx+cy = mx + c, giving mm and cc each to 3 significant figures. Use the formulas
m=nxyxynx2(x)2,c=ymxn.m = \frac{n\sum xy - \sum x\, \sum y}{n\sum x^2 - \left(\sum x\right)^2}, \qquad c = \frac{\sum y - m\sum x}{n}.
Hence predict the test score for a student who studies for 10 hours, giving your answer to the nearest integer. [5]
c
The outlier is removed. A student claims that removing this single point will increase the gradient of the line of best fit by more than 100%. Recalculate mm and cc for the remaining 7 data points (each to 3 significant figures) and justify whether the student's claim is correct. [5]
Question diagram

Solutions

38QuestionQuartiles and percentiles (discrete and continuous data)Assessment Practice
11 marks~17 minCriterion B
The table below shows five data sets, each containing 11 values listed in ascending order.

Data set A: 1,3,5,7,9,11,13,15,17,19,211, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21
Data set B: 1,3,5,7,9,11,13,15,17,19,611, 3, 5, 7, 9, 11, 13, 15, 17, 19, 61
Data set C: 1,3,5,7,9,11,13,15,17,19,811, 3, 5, 7, 9, 11, 13, 15, 17, 19, 81
Data set D: 1,3,5,7,9,11,13,15,17,41,611, 3, 5, 7, 9, 11, 13, 15, 17, 41, 61
Data set E: 1,3,5,7,9,11,13,15,37,41,611, 3, 5, 7, 9, 11, 13, 15, 37, 41, 61
a
Find the median and interquartile range (IQR) for data set A. [3]
b
Complete the summary below by finding the median, Q1Q_1, Q3Q_3, and IQR for data sets B, C, D, and E. All values must be exact.

Data set A — Median1111Q1Q_1: 55Q3Q_3: 1717IQR: 1212
Data set B — Median\_\_\_Q1Q_1: \_\_\_Q3Q_3: \_\_\_IQR: \_\_\_
Data set C — Median\_\_\_Q1Q_1: \_\_\_Q3Q_3: \_\_\_IQR: \_\_\_
Data set D — Median\_\_\_Q1Q_1: \_\_\_Q3Q_3: \_\_\_IQR: \_\_\_
Data set E — Median\_\_\_Q1Q_1: \_\_\_Q3Q_3: \_\_\_IQR: \_\_\_


Describe the pattern you observe as the large value(s) move further into the data set from the maximum position. [4]
c
A classmate claims: "Making any value in the upper half of an ordered data set larger will always increase Q3Q_3." Justify whether this claim is correct or incorrect, referring to the definition of quartiles as positional measures and using at least two data sets from part (b) as evidence. [4]
Question diagram

Solutions

39QuestionQuartiles and percentiles (discrete and continuous data)Assessment Practice
10 marks~15 minCriterion C
The cumulative frequency graph shows the waiting times, in minutes, for 80 patients at a health clinic. The graph passes through the points (0,0)(0, 0), (10,8)(10, 8), (20,24)(20, 24), (30,52)(30, 52), (40,72)(40, 72), and (50,80)(50, 80).

A student records the following attempt to find the quartiles:

Q1Q_1: the 20th patient waited 12 minutes.
Q2Q_2: the 40th patient waited 25 minutes.
Q3Q_3: the 60th patient waited 33 minutes.
a
State two errors in the student's method. [2]
b
Using correct statistical notation, explain how to read the lower quartile Q1Q_1 from a cumulative frequency graph for n=80n = 80 data values. [3]
c
Using the graph, determine the correct values of Q1Q_1, Q2Q_2, and Q3Q_3, and hence calculate the interquartile range. The clinic defines a "short wait" as any waiting time below the lower quartile. Justify whether a patient who waited 17 minutes experienced a short wait, giving your answer to the nearest minute. [5]
Question diagram

Solutions

40QuestionQuartiles and percentiles (discrete and continuous data)Assessment Practice
8 marks~12 minCriterion D
A school records the predicted IB Mathematics scores (on a scale of 1 to 7) for its 60 MYP 5 Extended students. The results are summarised in the cumulative frequency table below.

Score34567
Cumulative frequency416385660
a
Calculate the percentile rank of a student whose predicted score is 6, giving your answer to one decimal place. [2]
b
A university requires applicants to have a predicted score at or above the 75th percentile of their school cohort. A second student scored 5. Determine, showing all working, whether this student meets the university's requirement. [3]
c
A third student argues that achieving a high percentile rank within this school guarantees a strong university application. Critique this argument, referring to at least two distinct mathematical or statistical reasons. [3]
Question diagram

Solutions