RevisionPrep
Back to Blog

Statistics – Averages and Analysis

Grouped data, missing values, box plots and spotting misleading stats for IB MYP 3 Maths

Statistics dashboard showing a grouped frequency table, a box plot and a bar chart side by side
Subject
Mathematics
Curriculum
IB MYP
Grade
MYP 3
Topic
Statistics – Averages and Analysis
Reading
7 min
Difficulty
Standard

Quick facts

Difficulty
★★★☆☆
Assessed under
Criterion A & Criterion D
Prerequisites
Basic mean, median, mode, ordering data
You'll learn
Grouped data, estimated mean, missing values, IQR, misuse of stats
Revision time
35-45 min

Statistics questions in IB MYP 3 Maths rarely just ask you to calculate a mean — they ask you to reason about it. You'll organise raw data into grouped tables, work backwards from a given average to find a missing value, summarise spread with box plots and interquartile range, and judge whether a statistic is actually fair. The formulas themselves — midpoint, estimated mean, IQR — are simple tools; the real marks come from explaining what your number means and where it can go wrong. This teaser walks through the five ideas examiners test most under Criterion A and Criterion D: grouping data, estimating a mean from a table, reverse-engineering missing values, reading box plots correctly, and spotting how averages and graphs can mislead. For full worked examples, trap-avoidance drills and mock exam-style questions, the complete revision note is linked below.

What you’ll be able to do

Group raw data into class intervals of equal width
Calculate the midpoint of a class interval
Estimate the mean of grouped data using midpoints and frequencies
Find a missing value or frequency given the mean
Sort data and locate the five-number summary for a box plot
Calculate the interquartile range and explain what it shows
Compare two data sets using both centre and spread
Identify how averages and graphs can misrepresent data
1

Grouping Data: Trading Precision for Readability

When you have hundreds of raw values, grouping them into class intervals (like 30–34, 35–39...) makes the shape of the data visible at a glance. But grouping is a one-way trade: once values are grouped, the exact originals are gone, and every calculation from that table afterwards is an estimate. All class intervals in one table must share the same width so histogram bars stay fairly comparable.

Grouped frequency table with class intervals, midpoints and frequency column
Data typeExample intervalMeaningBoundary value goes in
Continuous (height, mass)140 ≤ h < 150Non-overlapping inequalityThe lower interval only
Discrete (test scores, ages)10 ≤ x ≤ 19Inclusive integersNext interval starts at 20

Exam tip

If asked what a histogram bar's height represents, use the word 'frequency' explicitly — 'the number of students whose score falls in that interval.' Vague answers like 'it shows the scores' lose the mark.

Common mistake

Using one giant interval (e.g. ages 6–18 in one group) hides the shape entirely — it's technically grouped but tells you nothing useful.

Mini summary

Grouping reveals shape but loses exact values — keep class widths equal.

2

Estimating the Mean from a Grouped Table

Because grouped data replaces every value in an interval with its midpoint, any mean calculated from a grouped table is only ever an estimate of the true mean. The method is to multiply each midpoint by its frequency, sum those products, then divide by the total frequency. Always write the midpoint calculation as its own line before multiplying, so the method mark is locked in even if a later arithmetic slip happens.

Worked calculation showing frequency times midpoint for each class interval summing to an estimated mean

Exam tip

If a question asks you to comment on whether the estimated total equals the real total, say explicitly that it's an estimate — because every value in an interval was treated as if it equalled the midpoint.

Common mistake

Multiplying frequency by the lower bound (or upper bound) of an interval instead of the midpoint — this single slip throws off the whole calculation.

Mini summary

Estimated mean = Σ(frequency × midpoint) ÷ Σfrequency — never exact once data is grouped.

3

Working Backwards: Finding a Missing Value from the Mean

Since mean = sum ÷ n, rearranging gives sum = mean × n — and this lets you solve for a missing score if you already know the mean and every other value. Always state n first by counting the TOTAL number of values (including the unknown one), then subtract the sum of the known values from the required total sum. The exact same logic applies to a missing frequency in a grouped table.

Diagram showing five test scores with one missing, and the equation sum equals mean times n rearranged to solve for it

Exam tip

Write 'n = ...' as your very first line of working, counted directly from the question, before touching any other numbers.

Common mistake

Multiplying the mean by the number of KNOWN values instead of the TOTAL number of values including the missing one — this is the single most common slip in this type of question.

Mini summary

Rearrange mean = sum ÷ n into sum = mean × n, then subtract the known sum from the total required sum.

4

Box Plots and Interquartile Range

A box plot compresses a whole ordered data set into five landmark numbers: minimum, lower quartile (Q1), median, upper quartile (Q3), and maximum. Data must always be sorted first, since every quartile is defined by position. Q1 is the median of the lower half and Q3 is the median of the upper half — for an odd number of values, exclude the overall median from both halves before splitting.

Box plot on a number line labelled minimum, lower quartile, median, upper quartile, maximum with IQR bracket
StatisticWhat it shows
MedianSplits the ordered data exactly in half
Box (Q1 to Q3)The middle 50% of the data — width shows spread
WhiskersRange of the outer 25% on each side
IQR = Q3 − Q1Spread of the middle 50%, ignoring extreme outliers

Common mistake

Including the overall median inside the lower or upper half when finding Q1 or Q3 — for odd n, this shifts both quartiles and gives the wrong IQR.

Mini summary

Sort first, split around the median, then IQR = Q3 − Q1 describes the spread of the middle half.

5

Spotting Misleading Statistics

Every grouping decision, choice of average, and axis scale is a decision that can hide as much as it reveals. Cramming 200 students aged 6–18 into one giant interval erases all sense of whether the group is mostly children or teenagers. The mean is dragged by extreme values while the median resists this, so reporting 'the average' without saying which one can mislead — and a bar chart with a truncated, non-zero axis can make a tiny difference look huge.

Two bar charts side by side, one with a zero-start axis and one with a truncated axis showing the same data looking different

Exam tip

When comparing two data sets, always quote a measure of centre (mean or median) AND a measure of spread (range or IQR) in the same sentence — one number alone is never a fair comparison.

Common mistake

Reporting small sample results as if representative, e.g. '3 out of 4 students prefer...', without acknowledging how few people were actually surveyed.

Mini summary

Question which average is used, whether intervals hide the shape, and whether axes are drawn fairly.

Quick formula sheet

Midpoint of a class interval, used to represent every value inside it.Add the two ends, halve it — like finding the middle of a road.
Estimated mean of grouped data using frequency and midpoint.Frequency × midpoint, add them all up, divide by total frequency.
Total sum of all data values equals the mean multiplied by the number of values — used to reverse-engineer missing values.Flip mean = sum ÷ n around: sum = mean × n.
Interquartile range — the spread of the middle 50% of the data.Top quarter cut minus bottom quarter cut = the 'middle chunk' width.

Practice questions

Easy
  1. Find the midpoint of the class interval 20–30.
  2. The mean of four numbers is 8. What is the sum of the four numbers?
  3. Order the data set 5, 9, 3, 7, 1 and state the median.
Medium
  1. A grouped table has intervals 0–10 (f=3), 10–20 (f=5), 20–30 (f=2). Estimate the mean.
  2. The mean of six scores is 12. Five of the scores are 8, 10, 14, 15 and 11. Find the sixth score.
  3. For the data set 4, 8, 9, 12, 15, 18, 21, find Q1, Q3 and the IQR.
Challenge
  1. Explain why grouping 200 ages into a single 6–18 interval hides important information, and propose better class intervals with a justification.
  2. A grouped frequency table has a missing frequency in one interval. Given the overall mean, set up and solve an equation to find it.
  3. Two classes have the same median test score but very different IQRs. Explain what this tells you about each class's performance, using both centre and spread.

Frequently asked questions

Why is the mean from grouped data called an 'estimate'?+

Because grouping replaces every individual value inside an interval with its midpoint, so the exact original numbers are lost and any mean calculated afterwards is only an approximation of the true mean.

How do I find a missing value if I only know the mean?+

Rearrange mean = sum ÷ n into sum = mean × n, count the total number of values including the missing one, then subtract the sum of the known values from that total.

What's the difference between range and interquartile range?+

Range uses the maximum and minimum, so it's affected by extreme outliers. IQR only uses Q1 and Q3, so it describes the spread of the middle 50% and ignores extreme values.

Do I need to sort data before finding quartiles?+

Yes — quartiles are defined by position in an ordered list, so unsorted data gives meaningless quartile values.

Why must class widths stay the same across a grouped table?+

Equal class widths keep histogram bars fairly comparable; unequal widths distort the shape and can mislead anyone reading the chart.

How can averages be used to mislead people?+

The mean can be pulled by extreme values while the median stays stable, so quoting 'the average' without saying which type was used can accidentally or deliberately misrepresent the data.

Get the Full MYP 3 Statistics Revision Notes

Step-by-step worked examples for grouped data, missing values, and box plots All common mistakes and examiner traps explained in full detail Original mock papers and exam-style questions to test yourself Clear breakdown mapped to Criterion A and Criterion D expectations
Get the Statistics – Averages and Analysis notes on RevisionPrep

Related articles