A cumulative frequency curve is built by plotting the running total of frequencies against the upper boundary of each class interval, then joining the points with a smooth curve. From this curve you read off the median (at 50% of total frequency), lower quartile (25%), and upper quartile (75%) by drawing horizontal lines across and dropping vertical lines down to the x-axis. These same five values — minimum, lower quartile, median, upper quartile, maximum — form a box-and-whisker plot.
This topic isn't conceptually hard, but it's mechanically unforgiving. Every value plotted against the wrong boundary, every reading taken from the wrong axis, and every quartile calculated with the wrong fraction cascades into a wrong final answer — even when the student understood exactly what a median or quartile represents.
The two things that separate full marks from partial marks are precision in constructing the table (upper boundaries, not midpoints) and precision in reading the graph (draw the lines, don't estimate by eye).
The single most common error happens here, before any graph is even drawn.
The rule that trips up most students
Cumulative frequency is always plotted against the upper class boundary — never the midpoint, and never the lower boundary. If a class interval is "10 ≤ x < 20", you plot against 20, not 15.
Example — test scores of 60 students
| Score (x) | Frequency | Upper boundary | Cumulative frequency |
|---|---|---|---|
| 0 ≤ x < 10 | 4 | 10 | 4 |
| 10 ≤ x < 20 | 9 | 20 | 13 |
| 20 ≤ x < 30 | 14 | 30 | 27 |
| 30 ≤ x < 40 | 18 | 40 | 45 |
| 40 ≤ x < 50 | 11 | 50 | 56 |
| 50 ≤ x < 60 | 4 | 60 | 60 |
Plot each (upper boundary, cumulative frequency) point, then join them with a single smooth curve — not straight lines, and not a curve that dips or wobbles. The curve should always be non-decreasing, since cumulative frequency can never go down.
For a data set of size n, use these positions on the cumulative frequency axis (if you have not yet covered mean, median, and mode from a frequency table, start there first):
Lower quartile (LQ)
n/4 th value
Median
n/2 th value
Upper quartile (UQ)
3n/4 th value
Interquartile range
UQ − LQ
Continuing the example — n = 60
60/2 = 30. Draw a horizontal line from 30 on the y-axis to the curve, then drop down to the x-axis. Read off: median ≈ 32.60/4 = 15. Draw across from 15, down to the x-axis. Read off: LQ ≈ 24.3 × 60/4 = 45. Draw across from 45, down to the x-axis. Read off: UQ ≈ 39.IQR = UQ − LQ = 39 − 24 = 15.Common mistake
Using n instead of n/2, n/4, and 3n/4 — or forgetting to actually draw the horizontal and vertical construction lines on the graph. Examiners expect to see these lines; a numerical answer alone without the lines drawn on the curve can lose marks even if correct, since the method must be shown on the graph itself.
A box-and-whisker plot displays five values on a single scale: minimum, lower quartile, median, upper quartile, and maximum. It's a compact visual summary of the same data your cumulative frequency curve already gave you.
Reading a box plot — the parts and what they mean
O-Level questions frequently show two box plots side by side (e.g. Class A vs Class B test scores) and ask you to compare them, which is the same skill as comparing two data sets more generally. There are two things to comment on, and both are required for full marks:
Always comment on both — students often compare only the medians and forget the spread, which usually costs half the marks on a comparison question. A complete answer always addresses average and consistency.
Why do we plot against the upper boundary and not the midpoint?
Cumulative frequency represents "the number of values less than or equal to this point." At the upper boundary of a class, you've accounted for every value in that class, so the cumulative total is only accurate at that boundary — not at the midpoint, which is partway through the class.
Is the cumulative frequency curve always S-shaped?
Typically yes, for data that's roughly normally distributed — this S-shape is called an ogive. It starts flat, rises steeply through the middle where most data is concentrated, then flattens again near the maximum. It should never decrease, since cumulative frequency can only increase or stay the same.
Do I use n/2 or (n+1)/2 for the median position on a cumulative frequency graph?
For a cumulative frequency curve with continuous/grouped data, always use n/2 (not (n+1)/2). The (n+1)/2 formula is used for finding the median directly from a small, ungrouped, listed data set — a different technique. Grouped/graphical cumulative frequency questions in O-Level use n/2, n/4, and 3n/4 throughout.
What's the difference between range and interquartile range?
Range = maximum − minimum, and includes all data including outliers. Interquartile range = upper quartile − lower quartile, and only reflects the spread of the middle 50% of data, making it less affected by extreme values. Examiners often ask which measure is "more appropriate" when outliers are present — the answer is IQR.
— Mr Gan Math Tuition
Mr. Gan works with students who want precise, mark-scheme-aligned technique — not just conceptual understanding.
Chat with Mr. Gan