CHAPTER 1
Analysis of Data
AQA Mathematical Studies · Paper 1
1.1 Data Types & Sampling
1.2 Averages & Stem-and-Leaf
1.3 Measures of Spread
1.4 Box & Whisker Plots
1.5 Cumulative Frequency
1.6 Histograms
AQA LEVEL 3 CERTIFICATE
Mathematical Studies · Paper 1
AQA MATHEMATICAL STUDIES — CHAPTER 1
Chapter Overview & Learning Objectives
SECTION 1.1
Data Types & Sampling
Qualitative, discrete & continuous data; random, stratified, cluster and quota sampling methods.
SECTION 1.2
Averages & Stem-and-Leaf
Mean, mode and median from tables; back-to-back stem-and-leaf diagrams for comparison.
SECTION 1.3
Measures of Spread
Range, interquartile range and standard deviation; identifying outliers using the 1.5 × IQR rule.
SECTION 1.4
Box & Whisker Plots
Five-number summary (min, Q1, median, Q3, max); plotting outliers; comparing distributions.
SECTION 1.5
Cumulative Frequency
Drawing CF graphs; reading percentiles; estimating median and IQR directly from the graph.
SECTION 1.6
Histograms
Frequency density on the y-axis; handling unequal class widths; interpreting histogram area.
Prerequisites: Frequency tables, pie charts, bar charts, and percentage calculations are assumed knowledge for this chapter.
1.1 Data Types & Sampling Methods
AQA Mathematical Studies — Chapter 1 · Paper 1 Topic
Data Types
THREE CATEGORIES
Qualitative (Categorical)
Descriptive — not a number
colour
gender
exam grade
nationality
Quantitative — Discrete
Exact countable values only
no. of texts
shoe size
ticket prices
age
Quantitative — Continuous
Any value within a range
height
weight
time
leaf area
Quick test: "Can it take any value in a range?" → Continuous. "Only exact/countable values?"→ Discrete. "A description?" → Qualitative.
Sampling Methods
Population — all members of the group being studied
Sample — sub-group used when population too large or costly
FOUR METHODS
Random
Every member has an equal chance of selection — unbiased but can be expensive
Stratified
Groups sampled in same proportion as population — most representative
Stratified Formula:
Sample from group = (group size ÷ total) × sample size
Cluster
Randomly select natural sub-groups (e.g. Local Authorities), then survey all within
Quota
Interviewers fill set quotas of types — quick but may introduce bias
Worked Example — Stratified Sample of 20 from 1522 students
Female Part time (465):
465÷1522×20 =
6
Female FullT (236):
236÷1522×20 =
3
Male Part Time (624):
624÷1522×20 =
8
Male FullT (197):
197÷1522×20 =
3
Check: 6 + 3 + 8 + 3 = 20 ✓ (must equal total sample size)
Stratified Sampling — Worked Examples
FORMULA
Sample from group = (Group size ÷ Total population) × Sample size
Round to nearest whole number · Check totals sum correctly
Key Rule: Always round each group to the nearest whole number, then check all group samples add up to the total sample size before finalising.
1
College Sample
Total = 1522 · n = 20
1522 students
Sample size: 20
4 groups
Female Part-Time
465 ÷ 1522 × 20 =
6
Female Full-Time
236 ÷ 1522 × 20 =
3
Male Part-Time
624 ÷ 1522 × 20 =
8
Male Full-Time
197 ÷ 1522 × 20 =
3
✓
Verification: 6 + 3 + 8 + 3 = 20
Matches sample size ✓
2
Company Sample
Total = 922 · n = 200
Department | M FT | M PT | F FT | F PT | Total |
Factory | 239 | 171 | 179 | 91 | 680 |
Warehouse | 55 | 20 | 28 | 15 | 118 |
Office | 29 | 7 | 42 | 5 | 83 |
Delivery | 17 | 4 | 1 | 20 | 42 |
922 employees
Sample size: 200
4 departments
DEPARTMENT BREAKDOWN
WORKED CALCULATION — FACTORY MALE FT
239 ÷ 922 × 200 = ≈ 52
Apply the same formula to every group. All results must sum to 200.
✓
All group samples must sum to 200
Always verify totals ✓
Exercise 1A — Skills
SKILL GUIDE
DATA TYPE CLASSIFICATION
Qualitative
Descriptive — not a number
colour, gender, nationality, grade
Discrete
Countable exact values only
number of texts, shoe size, ticket prices
Continuous
Any value in a range (measured)
weight, height, time, length, area
STRATIFIED SAMPLING — METHOD
1
Find the total population size
2
Apply the formula for each group:
Sample from group = (Group size ÷ Total) × n
3
Round each result to the nearest whole number
4
Check: all group samples must sum to total sample size
Group | Group Size | Total | Sample Size (n) | Calculation | Result |
Year 12 | 90 | 240 | 40 | (90 ÷ 240) × 40 | = 15 |
Year 13 | 150 | 240 | 40 | (150 ÷ 240) × 40 | = 25 |
WORKED EXAMPLE — STRATIFIED SAMPLING
Example: A school has 240 students. Year 12 = 90, Year 13 = 150. Take a stratified sample of 40.
Verification: 15 + 25 = 40 ✓ (totals match sample size)
SURVEY METHODS — KEY POINTS
Random
Every member has equal chance of selection. Unbiased but can be expensive and time-consuming.
Stratified
Groups sampled in proportion to population. Representative of all subgroups. Use the formula.
Cluster
Randomly select natural sub-groups (e.g. schools, regions), then survey all within chosen clusters.
Quota
Interviewers fill set quotas by type. Quick and practical but may introduce selection bias.
Exam Tip: Always show your calculation clearly and verify the total. If rounding causes the sum to be off by 1, adjust the largest group. State which method you used and why it is appropriate.
Exercise 1A — Questions
ATTEMPT ALL QUESTIONS
Work through all six questions before checking answers on the next slide.
Q1
DATA TYPES
Number of texts sent
Colour of flowers
Weight of puppies
Exam grade
Shoe size
Length of feet
Time asleep
Leaf area
Ticket prices
Gender of fans
Age of fans
Nationality of players
Classify each item as Qualitative , Discrete or Continuous :
Q2
DATA SOURCES
(a) Give one example of a primary data source you could use.
(b) Give one example of a secondary data source you could use.
Describe how you would find both primary and secondary data on students' earnings from part-time work.
Q3
SURVEY METHODS
(a) Suggest two questions the company might ask.
(b) Comment on three methods: bus station interviews , postal questionnaires , and radio/TV invitation to respond .
A market research company surveys people about a new cinema complex .
Q4
STRATIFIED SAMPLING — COMPANY SURVEY
Formula: (Group size ÷ 922) × 200
(a) Calculate the sample size for Factory Male FT . Show your working.
(b) Calculate the sample sizes for all other groups .
(c) Verify your answers sum to 200. ✓
A company wants a stratified sample of 200 employees from its 922 staff . The two-way table shows the breakdown by department and employment type.
Department | Male FT | Male PT | Female FT | Female PT | Total |
Factory | 239 | 171 | 179 | 91 | 680 |
Warehouse | 55 | 20 | 28 | 15 | 118 |
Office | 29 | 7 | 42 | 5 | 83 |
Delivery | 17 | 4 | 1 | 20 | 42 |
TOTAL | 340 | 202 | 250 | 131 | 922 |
Q5
QUOTA SAMPLING
(a) State what quotas you would set and why.
(b) Explain how interviewers would fill each quota.
(c) Give one advantage and one disadvantage of this method.
Describe how quota sampling could be used to survey 500 people about a proposed new airport .
Q6
CLUSTER SAMPLING
Describe how cluster sampling could be used to investigate health care provision across England .
Exercise 1A — Answers
AQA CHAPTER 1.1
Q1
DATA TYPE CLASSIFICATION
Qualitative
Discrete
Continuous
Colour of flowers
No. of texts sent
Weight of puppies
Exam grade
Shoe size
Length of feet
Gender of fans
Ticket prices
Time asleep
Nationality of players
Age of football fans
Leaf area
Qualitative = descriptions | Discrete = countable exact values | Continuous = any value in a range
Q4
STRATIFIED SAMPLING — 1200 STUDENTS, SAMPLE OF 50
Formula: (Group Size ÷ Total) × Sample Size
Year 12 (480):
(480 ÷ 1200) × 50 =
20 students
Year 13 (720):
(720 ÷ 1200) × 50 =
30 students
Verification: 20 + 30 = 50 ✓
Selection: use random numbers or names in a hat within each year group
Q2
DATA SOURCES — STUDENT EARNINGS
PRIMARY DATA
Conduct your own survey or questionnaire asking students directly about their part-time earnings.
SECONDARY DATA
Use existing records from HMRC, college databases, or government statistics on student employment.
Q3
SURVEY METHOD CRITIQUE — CINEMA COMPLEX
Bus Station
Biased — only reaches evening/Saturday users; not representative of all potential cinema-goers.
Postal
Low response rate — may not reach all areas or demographics equally; self-selection by motivated respondents.
Radio/TV
Self-selection bias — only motivated or opinionated people respond; not representative of general public.
Q5
QUOTA SAMPLING
Q6
CLUSTER SAMPLING
Q5 — QUOTA SAMPLING (AIRPORT)
Set quotas by age and gender in proportion to the local population. Interviewers select participants until each quota is filled.
Not random — interviewers choose who to approach.
Q6 — CLUSTER SAMPLING (HEALTH CARE)
Randomly select a sample of Local Authorities (clusters), then survey all people within those LAs about health care provision.
Not a random sample of England's entire population.
1.2 Averages: Mean, Mode & Median
AQA CHAPTER 1.2
WORKED EXAMPLE — DRIVING TEST AGES
GROUP
MODE
MEAN
MEDIAN
NOTE
Female
17
19
Outlier age 49 inflates mean
Male
18
19.9
19.5
All three averages close together
Key Insight: The female mean (22.7) is distorted by the outlier age 49 — the median (19) better represents the typical female candidate. When outliers are present, the median is more reliable than the mean.
Mode & Median
Mode — easy to find; always a data value; may not exist or may be bimodal
Mode — does not use all data values in its calculation
Median — often a data value; not affected by extreme values (outliers)
Median — does not use all data; robust to skewed distributions
Mean
x̄ = Σx ÷ n
Uses all data values — every value contributes to the result
Affected by extreme values — outliers pull the mean up or down significantly
Involves more calculation — sum all values then divide by n
Not usually a data value — result is often a decimal
22.7
⚠ distorted
Example 1 & 2: Stem-and-Leaf Diagrams
AQA Chapter 1.2 — Constructing and interpreting stem-and-leaf diagrams
Example 1 — Test Marks (out of 60)
DATA: 39, 37, 56, 44, 32, 40, 26, 58, 31, 42, 37, 51, 29, 38, 28, 42 (N = 16)
STEM | LEAF (ORDERED)
2
6 8 9
3
1 2 7 7 8 9
4
0 2 2 4
5
1 6 8
Key: 5 | 1 means 51 marks · n = 16
Mode
37
and
42
— bimodal (each appears twice)
Median
= (8th + 9th) ÷ 2 = (38 + 39) ÷ 2 =
38.5 marks
n
16 values → middle pair: 8th & 9th
Example 2 — Driving Test Ages (Back-to-Back)
◀ Female (leaves)
Stem
(leaves) Male ▶
9 9 8 7 7 7
1
7 7 8 8 8 8 9
6 5 2 1
2
0 0 1 2 2 3 6
9
4
—
Key: 5 | 2 | 6 means Female 25 , Male 26
Female
Median =
19
· Range =
32
(outlier: age 49)
Male
Median =
19.5
· Range =
9
Female ages are far more widely spread than male ages — the outlier age 49 inflates
the female range to 32, compared to just 9 for males. Both groups share a similar median
(~19).
Always include a Key
— it must be shown on every stem-and-leaf diagram.
AQA MATHEMATICAL STUDIES — CHAPTER 1.3
Measures of Spread: Range, IQR & Standard Deviation
Range = Max − Min
IQR = UQ − LQ
Use σₙ₋₁ on calculator
Key Insight: Both classes share the same median (59) , but Class B is far more spread out— IQR = 33 vs IQR = 7 for Class A. Same centre, very different spread.
Standard Deviation: Use the σₙ₋₁ button on your calculator (or STDEV in Excel). A higher σmeans the data is more spread out from the mean.
Measure | Class A (n=15) | Class B (n=14) |
Lower Quartile (LQ) | 56 | 37 |
Median | 59 | 59 |
Upper Quartile (UQ) | 63 | 70 |
IQR = UQ − LQ | 7 | 33 |
Spread | Compact | Much more spread |
CLASS A VS CLASS B — COMPARISON (N=15 AND N=14)
OUTLIER RULE & WORKED EXAMPLE
OUTLIER DEFINITION
Value < LQ − 1.5 × IQR
Value > UQ + 1.5 × IQR
CLASS A — OUTLIER CHECK
LQ − 1.5 × IQR = 56 − 1.5 × 7
= 56 − 10.5 = 45.5
Value 8 < 45.5
Value 8 IS an outlier ✓
RANGE
Easy to find. Affected by extreme values (outliers). Uses only two data points.
IQR (INTERQUARTILE RANGE)
Uses middle 50% of data. Not affected by outliers. For n ≤ 20: split data either side of median.
STANDARD DEVIATION Σ
Uses ALL data. Gives fair comparison. More complex to calculate. Use STDEV in spreadsheet.
IQR Worked Examples 1 & 2
AQA CH 1.3
Key Insight: IQR is more robust than range — it ignores extreme values. Use the rule: outlier if value < LQ − 1.5×IQR or value > UQ + 1.5×IQR.
EXAMPLE 1
Class A Test Scores
n = 15
SORTED DATA
8 , 54, 55, 56, 58, 59, 59, 59 , 62, 62, 63, 63, 6 3, 64, 65
Median
8th value = 59
Lower half
8, 54, 55, 56, 58, 59, 59
LQ (4th)
= 56
Upper half
62, 62, 63, 63, 63, 64, 65
UQ (4th)
= 63
IQR
63 − 56
= 7
OUTLIER CHECK
LQ − 1.5 × IQR = 56 − 10.5 = 45.5
Value 8 < 45.5
8 IS an outlier ✓
EXAMPLE 2A
Group P
n = 11
SORTED DATA
18, 19, 19, 21, 23, 23 , 26, 27, 32, 35, 43
Median
6th value = 23
Lower half
18, 19, 19, 21, 23
LQ (3rd)
= 19
Upper half
26, 27, 32, 35, 43
UQ (3rd)
= 32
IQR
32 − 19
= 13
OUTLIER CHECK
LQ − 1.5 × IQR = 19 − 19.5 = −0.5
UQ + 1.5 × IQR = 32 + 19.5 = 51.5
All values within bounds
No outliers
EXAMPLE 2B
Group Q
n = 10
SORTED DATA
20, 21, 24, 25, 29, 31, 33, 35, 37, 69
Median
(29 + 31) ÷ 2 = 30
Lower half
20, 21, 24, 25, 29
LQ (3rd)
= 24
Upper half
31, 33, 35, 37, 69
UQ (3rd)
= 35
IQR
35 − 24
= 11
OUTLIER CHECK
UQ + 1.5 × IQR = 35 + 16.5 = 51.5
Value 69 > 51.5
69 IS an outlier ✓
Exercise 1B — Skills
HOW TO FIND MEAN, IQR & STANDARD DEVIATION
FINDING THE MEAN
1
Add all values together: Σx
2
Divide by the number of values:
x̄ = Σx / n
Frequency table: multiply each value by its frequency, sum, then divide by total frequency
x̄ = Σfx / Σf
Weighted mean (combining groups):
(n₁x̄₁ + n₂x̄₂) / (n₁ + n₂)
Do NOT simply average the two means!
FINDING THE IQR
1
Sort data in ascending order
2
Find the median (middle value)
3
For n ≤ 20 : split data either side of median (exclude median itself)
4
LQ = median of lower half | UQ = median of upper half
5
IQR = UQ − LQ
Outlier rule:
value < LQ − 1.5×IQR or value > UQ + 1.5×IQR
STANDARD DEVIATION
Use the
σₙ₋₁
button on your calculator
In a spreadsheet: use the STDEV function
Higher σ = data is more spread out from the mean
Standard deviation uses ALL data values — gives a fair comparison between datasets
Use σ to compare spread when two datasets have the same mean but different variability
TARGET MEAN — WORKED EXAMPLE
Formula: Runs needed = (target mean × new n) − current total
Runs needed = (target x̄ × new n) − (current x̄ × old n)
JACK'S CRICKET SCORES
Jack's mean from 8 matches = 45
Target mean for 9 matches = 50
Current total = 8 × 45 = 360
Target total = 9 × 50 = 450
Runs needed in match 9:
450 − 360 = 90 runs
Always find both totals first, then subtract to find the missing value
Exercise 1B — Questions
ATTEMPT BEFORE CHECKING ANSWERS
Q1 — AVERAGES & RANGE
Find the mode, median, mean and range for the 20 Premiership season ticket prices below.
£1014
£335
£499
£595
£550
£544
£501
£365
£710
£299
£532
£383
£499
£608
£459
£400
£449
£795
£349
£640
Group | Number | Wage/hr |
Packers | 42 | £9.50 |
Fork-lift drivers | 14 | £13.50 |
Q3 — WEIGHTED MEAN
Stacy claims the overall mean wage is £11.50/hr. Explain why she is wrong and calculate the correct mean.
Hint: Use weighted mean — do NOT simply average the two wages.
Q4 — TARGET MEAN
How many runs must Jack score in match 9 to achieve an overall mean of 50?
8
Matches played
→
45
Current mean
→
50
Target mean (9 matches)
→
?
Runs in match 9
Hint: Target total = target mean × new n. Runs needed = target total − current total.
Q5 — MEDIAN & IQR
For each dataset find the median and interquartile range (IQR).
(a) Bus late arrival times (minutes) — n = 15
0, 0, 0, 5, 5, 10, 10, 10, 15, 15, 20, 25, 30, 35, 40
(b) Seeds germinating — n = 15
3, 4, 4, 5, 5, 5, 6, 6, 7, 8, 9, 10, 11, 12, 15
(c) Daily lunch spending — n = 11
£2.50, £3.00, £3.50, £3.50, £4.00, £4.50, £5.00, £5.50, £6.00, £7.00, £8.00
Q6 — MEAN & STANDARD DEVIATION
Find the mean and standard deviation for each dataset. Use σn−1 on your calculator.
(a) Test marks — n = 10
35, 38, 40, 40, 42, 43, 44, 45, 46, 47
(b) Daily temperatures (°C) — n = 10
12, 14, 15, 16, 17, 18, 19, 20, 21, 23
Calculator tip: Enter data → STAT mode → use σ
n−1
(not σ
n
) for sample standard deviation.
Answers on the next slide →
Exercise 1B — Answers
AQA CHAPTER 1.3
Weighted mean formula:
(n₁x̄₁ + n₂x̄₂) ÷ (n₁+n₂)
Target mean:
runs needed = (target mean × new n) − current total
Standard deviation: use
σₙ₋₁
button on calculator
Q1 / AVERAGES & RANGE
Season Ticket Prices (n = 20)
MODE
£499
MEDIAN
£500
MEAN
£526.30
RANGE
£715
1
Sorted data: £299 … £1014. Mode = £499 (appears twice)
2
Median = (10th + 11th) ÷ 2 = (£499 + £501) ÷ 2 = £500
3
Mean = £10,526 ÷ 20 = £526.30 | Range = £1014 − £299 = £715
Q3 / WEIGHTED MEAN
Stacy's Error — 42 Packers & 14 Fork-lift Drivers
STACY'S (WRONG)
£11.50/hr
CORRECT MEAN
£10.50/hr
✗
Stacy averaged the two means without weighting by group size
✓
(42 × £9.50 + 14 × £13.50) ÷ 56 = (£399 + £189) ÷ 56 = £10.50/hr
Q4 / TARGET MEAN
Jack's Cricket Scores — Match 9 Target
RUNS NEEDED
90 runs
1
Current total = 8 × 45 = 360
2
Target total = 9 × 50 = 450
3
Runs needed = 450 − 360 = 90 runs
Q5(A) / MEDIAN & IQR
Bus Late Arrival Times (n = 15, minutes)
MEDIAN
10
LQ
5
UQ
25
IQR
20
1
Sorted: 0,0,0,5,5, 10 ,10,10,15,15,20,25,30,35,40. Median (8th) = 10
2
Lower half: 0,0,0,5,5,10,10 → LQ (4th) = 5 | Upper half → UQ (12th) = 25
3
IQR = 25 − 5 = 20
Q6(A) / MEAN & STANDARD DEVIATION
Test Marks: 35, 38, 40, 40, 42, 43, 44, 45, 46, 47
MEAN (X̄)
42
STD DEV (Σ)
≈ 3.3
1
Σx = 35+38+40+40+42+43+44+45+46+47 = 420
2
Mean = 420 ÷ 10 = 42 | Use σₙ₋₁ on calculator → σ ≈ 3.3
1.4 Box and Whisker Plots
EXAMPLE 3
Manchester United Goals — 2013–14 Season (n = 38 matches)
AQA Mathematical Studies — Chapter 1.4 Example 3 | Man Utd 2013–14 Season
BOX PLOT DIAGRAM
⚠ Whisker drawn to 6 (last non-outlier). Value 7 marked with × as outlier.
0
Min
0.5
LQ
2
Median
3
UQ
6
Max*
0
1
2
3
4
5
6
7
×
Min=0
LQ=0.5
Med=2
UQ=3
Max=6
Outlier
Goals per Match
FREQUENCY TABLE — GOALS PER MATCH
= (9×0 + 9×1 + 9×2 + 7×3 + 4×4 + 0×5 + 2×6 + 1×7) ÷ 38
= (0 + 9 + 18 + 21 + 16 + 0 + 12 + 7) ÷ 38 = 64 ÷ 38
x̄ = 1.68 goals
Goals (x) | Freq (f) | Cum. Freq | fx |
0 | 9 | 9 | 0 |
1 | 9 | 18 | 9 |
2 | 9 | 27 | 18 |
3 | 7 | 34 | 21 |
4 | 4 | 38 | 16 |
5 | 0 | 38 | 0 |
6 | 2 | 40 | 12 |
7 × | 1 | 41 — wait | 7 |
Total | 38 | — | 64 |
KEY STATISTICS
📊
MEDIAN (19TH VALUE)
2 goals
📦
LQ (9.5TH VALUE)
0.5 goals
📦
UQ (28.5TH VALUE)
3 goals
📐
MEAN X̄ = ΣFX/ΣF
1.68 goals
σ
STANDARD DEVIATION Σ
1.32 goals
IQR CALCULATION
IQR = UQ − LQ
= 3 − 0.5 = 2.5 goals
⚠
OUTLIER CHECK
Upper fence = UQ + 1.5 × IQR
= 3 + 1.5 × 2.5
= 3 + 3.75 = 6.75
Value 7 > 6.75
∴ 7 goals IS an outlier ✓
Mean Formula: x̄ = Σfx ÷ Σf
Box & Whisker Plots — Examples 1 & 2
AQA Chapter 1.4 · Comparing distributions using five-number summaries
Key Rule: Always compare BOTH a measure of average (median) AND a measure of spread (IQR) when comparing two distributions from box plots.
Statistic | Class A | Class B |
Min | 8outlier | 25 |
LQ | 56 | 37 |
Median | 59 | 59 |
UQ | 63 | 70 |
Max | 65 | 80 |
Example 1 — Class Marks (n = 15)
FIVE-NUMBER SUMMARY
Class A — IQR
7
Class B — IQR
33
Same median (59) — but Class B's IQR is 33 vs 7 : Class B marks are far more spread out. Class A has an outlier at 8 (below LQ − 1.5 × IQR = 45.5).
Example 2 — Cricket Scores (n = 11)
PLAYER COMPARISON
Ben
Min 0
LQ 6
Median 29
UQ 47
Max 91
Ben — IQR
41
Ben — Mean / SD
33.0 / 29.5
Sanjay
Min 14
LQ 25
Median 29
UQ 49
Max 61
Sanjay — IQR
24
Sanjay — Mean / SD
35.6 / 14.8
Pick
Sanjay
— higher mean (35.6 > 33.0) AND far more consistent (SD 14.8 vs 29.5)
SKILLS
Exercise 1C — How to Draw & Compare Box Plots
HOW TO DRAW A BOX PLOT
1
Sort data in ascending order
2
Find the five-number summary : Min, LQ, Median, UQ, Max
3
For n ≤ 20 — exclude median when finding LQ and UQ
4
Draw a number line with a suitable scale
5
Draw a box from LQ to UQ
6
Mark the median inside the box with a vertical line
7
Draw whiskers from box to min and max
8
Mark outliers as separate × points
OUTLIER FENCES
Lower fence =
LQ − 1.5 × IQR
Upper fence =
UQ + 1.5 × IQR
Any value outside these fences is an outlier
COMPARING TWO DISTRIBUTIONS
Always comment on median — measure of average
Always comment on IQR — measure of spread
WORKED EXAMPLE — BEN VS SANJAY (CRICKET SCORES, N = 11)
COMPARISON CONCLUSION
Median: Both players have the same median of 29 — similar average performance.
IQR: Sanjay IQR = 24 < Ben IQR = 41 — Sanjay is more consistent (less spread).
Always compare BOTH median (average) AND IQR (spread) in your answer
Min | LQ | Median | UQ | Max |
0 | 6 | 29 | 47 | 91 |
Ben
SORTED DATA
0, 5, 6, 12, 20, 29 , 34, 43, 47, 75, 91
LQ = 3rd value = 6 | UQ = 9th value = 47
IQR = 47 − 6 = 41
Outlier check: UQ + 1.5 × 41 = 47 + 61.5 = 108.5
Max = 91 < 108.5 → No outliers
Min | LQ | Median | UQ | Max |
14 | 25 | 29 | 49 | 61 |
Sanjay
SORTED DATA
14, 18, 25, 27, 28, 29 , 37, 46, 49, 58, 61
LQ = 3rd value = 25 | UQ = 9th value = 49
IQR = 49 − 25 = 24
Outlier check: UQ + 1.5 × 24 = 49 + 36 = 85
Max = 61 < 85 → No outliers
Exercise 1C — Questions
ATTEMPT ALL QUESTIONS BEFORE CHECKING ANSWERS
Age | 17 | 18 | 19 | 20 | 21 |
Female | 6 | 9 | 11 | 7 | 0 |
Male | 13 | 11 | 8 | 5 | 9 |
Q1 — APPRENTICE AGES
Compare distributions for female and male apprentices (ages 17–21)
Compare the modal age , range , median age and mean age for each group. What do the differences tell you?
Q2 — CRICKET SCORES
Ben vs Sanjay — draw box plots and compare
BEN'S SCORES (N=11)
75
43
12
6
0
20
34
47
29
5
91
SANJAY'S SCORES (N=11)
28
49
27
61
29
14
58
37
25
18
46
Draw two box plots on the same scale. Compare mean and standard deviation . Who would you pick for the team and why?
Q3 — SUPERMARKET WAGES
Find median, IQR and check for outliers
£21,000
× 20 workers
£24,000
× 70 workers
£70,000
× 1 worker
Find the median and IQR . Show £70,000 is an outlier using UQ + 1.5 × IQR . Comment on the manager's claim that the average wage is £23,500 .
Q4 — BURGLARY RATES
Compare Badley and Lootham — find mean, SD, median, IQR and draw box plots
BADLEY (BURGLARIES/MONTH)
8
12
15
17
18
20
22
25
28
32
35
40
LOOTHAM (BURGLARIES/MONTH)
14
16
18
19
20
21
22
23
24
25
26
28
Find mean, SD, median and IQR for each town. Draw box plots and compare. Use statistical evidence to argue which town is safer and why.
Exercise 1C — Answers
MODEL ANSWERS
Q1
Apprentice Ages — Comparing Distributions
FEMALE (17–21)
Mode: 19
Range: 4
Median: 19
Mean: 18.9
MALE (17–21)
Mode: 17
Range: 4
Median: 18
Mean: 18.5
CONCLUSION
Males are slightly younger on average — all three averages are lower for males. Both groups have equal spread (range = 4 for both).
Q2
Cricket Scores — Ben vs Sanjay
Player | Mean | Std Dev (σ) | Verdict |
Ben | 33.0 | 29.5 | High spread — inconsistent |
Sanjay | 35.6 | 14.8 | Higher mean, far more consistent |
PICK SANJAY
Higher mean score (35.6 vs 33.0) AND much lower standard deviation (14.8 vs 29.5) — better average and far more consistent performance.
Q3
Supermarket Wages — Outlier Check
FIVE-NUMBER SUMMARY
Median: £24,000
LQ: £21,000
UQ: £24,000
IQR: £3,000
OUTLIER FENCE CALCULATION
UQ + 1.5 × IQR = £24,000 + 1.5 × £3,000 = £28,500
Since £70,000 > £28,500 → £70,000 IS an outlier ✓
MANAGER'S CLAIM
Manager's £23,500 is the mean — distorted by the £70,000 outlier. Median £24,000 is more representative of typical wages.
Q4
Burglary Rates — Badley vs Lootham
Badley
Mean: 20.8
SD: 11.5
Median: 17
IQR: 22
Lootham
Mean: 20.3
SD: 9.0
Median: 20
IQR: 14
KYLIE — FEWER ACCIDENTS
Badley median = 17 < Lootham median = 20 ✓ Lower typical burglary rate.
WINSTON — MORE CONSISTENT
Lootham SD = 9.0 < Badley SD = 11.5 ✓ More predictable crime levels.
BOTH ARGUMENTS VALID
Depends on which measure you prioritise — median (average) or standard deviation (consistency). Always justify your choice with statistical evidence.
1D: Cumulative Frequency Graphs
Building and reading CF graphs to find median, quartiles and percentiles from grouped data
Key Rules
Plot at
upper class boundary
Median at
n ÷ 2
LQ at
n ÷ 4
UQ at
3n ÷ 4
Percentile:
p% × n
Always title & label axes!
Build
CF Table
Hours worked (x) | Frequency (f) | Upper boundary |
0 < x ≤ 4 | 3 | 4 |
4 < x ≤ 8 | 29 | 8 |
8 < x ≤ 12 | 43 | 12 |
12 < x ≤ 16 | 10 | 16 |
16 < x ≤ 20 | 6 | 20 |
20 < x ≤ 24 | 1 | 24 |
Total | 92 | — |
Step 1 — Frequency Table
Year 12 students' part-time hours per week ( n = 92 )
1
Add frequencies cumulatively row by row
2
Always plot CF against the
upper class boundary
3
Draw a smooth S-curve through the plotted points
Upper boundary (x ≤) | Cumulative Frequency | Calculation |
4 | 3 | 3 |
8 | 32 | 3 + 29 |
12 | 75 | 32 + 43 |
16 | 85 | 75 + 10 |
20 | 91 | 85 + 6 |
24 | 92 | 91 + 1 |
Step 2 — Cumulative Frequency Table
Plot each pair (upper boundary, CF) on your graph axes
Reading off the graph (n = 92)
Median: 46th value →
≈ 9.3 hrs
LQ: 23rd value →
≈ 7 hrs
UQ: 69th value →
≈ 11.5 hrs
IQR: 11.5 − 7 →
≈ 4.5 hrs
90th percentile: 90% × 92 = 82.8th →
≈ 15 hrs
Exercise 1D — Questions & Answers
AQA CHAPTER 1D
Mass m (kg) | Frequency | Cum. Freq. | Upper Boundary |
0 < m ≤ 100 | 11 | 11 | 100 |
100 < m ≤ 150 | 31 | 42 | 150 |
150 < m ≤ 200 | 16 | 58 | 200 |
200 < m ≤ 250 | 28 | 86 | 250 |
250 < m ≤ 300 | 19 | 105 | 300 |
300 < m ≤ 400 | 5 | 110 | 400 |
Total | 110 | — | — |
Q1
Strawberry Farmer — Daily Mass Picked (n = 110 days)
A farmer records the mass of strawberries picked each day over 110 days. Use the frequency table to draw a cumulative frequency graph and find key statistics.
TASKS
(a)
State the modal class.
(b)
Draw a cumulative frequency graph (plot at upper class boundary).
(c)
Find the median, LQ, UQ, IQR, 40th and 80th percentiles.
(d)
Calculate the mean and standard deviation using midpoints.
ANSWERS
(A) MODAL CLASS
200 < m ≤ 250
(C) MEDIAN (55TH VALUE)
≈ 165 kg
LQ (28TH VALUE)
≈ 130 kg
UQ (83RD VALUE)
≈ 230 kg
IQR
≈ 100 kg
40TH PERCENTILE (44TH)
≈ 155 kg
80TH PERCENTILE (88TH)
≈ 250 kg
(D) MEAN & SD
x̄ ≈ 185 kg, σ ≈ 72 kg
Attendance | Frequency | Cum. Freq. | Upper Boundary |
Under 10,000 | 1 | 1 | 10,000 |
10,000 – 20,000 | 2 | 3 | 20,000 |
20,000 – 30,000 | 10 | 13 | 30,000 |
30,000 – 40,000 | 19 | 32 | 40,000 |
40,000 – 60,000 | 5 | 37 | 60,000 |
60,000 – 80,000 | 1 | 38 | 80,000 |
Total | 38 | — | — |
Q2
Football Club Attendance — Season (n = 38 matches)
A football club records attendance at each of its 38 home matches. Use the grouped frequency table to draw a cumulative frequency graph and compare with last year's data (mean = 38,000).
TASKS
(a)
State the modal class.
(b)
Draw a cumulative frequency graph.
(c)
Find the median and IQR from your graph.
(d)
Estimate how many matches had attendance over 50,000. Is this more than 5% of matches?
(e)
Calculate the mean and standard deviation.
(f)
Compare this year's data with last year (mean = 38,000).
ANSWERS
(A) MODAL CLASS
30,000 – 40,000
(C) MEDIAN (19TH VALUE)
≈ 33,000
IQR
≈ 10,000
(D) OVER 50,000
≈ 2 matches
5% OF 38 = 1.9
Yes — just above 5%
(E) MEAN & SD
x̄ ≈ 33,500, σ ≈ 10,000
(f) Comparison: Mean this year (≈33,500) is lower than last year (38,000), but SD is similar — attendances are slightly lower on average but equally spread. The lower mean may reflect fewer high-attendance fixtures.
1E: Histograms — Key Concepts & Frequency Density
AQA Mathematical Studies · Chapter 1E · School Traffic Survey Example
Tom's Mistake: Tom used HEIGHT to represent frequency —this is misleading when class widths are unequal. Always use AREA (FD × width) to represent frequency.
When to use histograms: Use histograms for continuous data with unequal class widths. No gaps between bars. Plot FD on y-axis, class boundaries on x-axis.
Activity 16: Verify each bar's area equals its frequency. Estimate the number of vehicles travelling at less than 15 mph using the histogram.
Frequency Density = Frequency ÷ Class Width
Key Formula
Area = Frequency
NOT Height
Frequency = FD × Width
Reading a Histogram
Speed (mph) | Frequency | Class Width | Freq. Density | Area Check |
0 < x ≤ 10 | 1 | 10 | 0.1 | 0.1 × 10 = 1 ✓ |
10 < x ≤ 20 | 8 | 10 | 0.8 | 0.8 × 10 = 8 ✓ |
20 < x ≤ 26 | 6 | 6 | 1.0 | 1.0 × 6 = 6 ✓ |
26 < x ≤ 30 | 28 | 4 | 7.0 | 7.0 × 4 = 28 ✓ |
30 < x ≤ 32 | 14 | 2 | 7.0 | 7.0 × 2 = 14 ✓ |
32 < x ≤ 34 | 3 | 2 | 1.5 | 1.5 × 2 = 3 ✓ |
34 < x ≤ 36 | 0 | 2 | 0 | 0 × 2 = 0 ✓ |
36 < x ≤ 38 | 2 | 2 | 1.0 | 1.0 × 2 = 2 ✓ |
38 < x ≤ 40 | 1 | 2 | 0.5 | 0.5 × 2 = 1 ✓ |
Histogram: Frequency Density vs Speed — Total = 65 vehicles
Exercise 1E — Q1 & Q2
AQA CHAPTER 1E · HISTOGRAMS & FREQUENCY DENSITY
Key Formula:
Frequency Density (FD) = Frequency ÷ Class Width
Given
Complete this
FD value
Q1 — Supermarket Spending (£P)
TASKS
(a) Copy and complete the table — state any assumptions you make.
(b) Draw a histogram to represent this data.
* Assumption needed: P < 25 starts at 0; P ≥ 250 ends at 350 (width = 100)
Class (£P) | Freq. | LCB | UCB | Width | FD = f ÷ w |
P < 25 | 75 | ? | 25 | 25 | 3.00 |
25 ≤ P < 50 | 97 | 25 | 50 | ? | 3.88 |
50 ≤ P < 100 | 165 | 50 | 100 | 50 | ? |
100 ≤ P < 150 | 86 | 100 | 150 | ? | 1.72 |
150 ≤ P < 250 | 23 | 150 | 250 | 100 | ? |
P ≥ 250 | 8 | 250 | 350* | 100* | ? |
Q2 — UK Population by Age (June 2013)
TASKS
(a) Copy and complete the table — state any assumptions you make.
(b) Draw a histogram to represent the UK age distribution.
* Assumption needed: 90+ ends at 100 (width = 10) — state this clearly
Age Group | Freq. (000s) | LCB | UCB | Width | FD = f ÷ w |
0 – 4 | 4,014 | 0 | 5 | ? | 802.8 |
5 – 15 | 11,179 | 5 | 16 | 11 | ? |
16 – 44 | 21,453 | 16 | 45 | ? | 739.8 |
45 – 64 | 16,328 | 45 | 65 | 20 | ? |
65 – 74 | 6,031 | 65 | 75 | ? | 603.1 |
75 – 89 | 4,574 | 75 | 90 | 15 | ? |
90+ | 527 | 90 | 100* | 10* | ? |
· For open-ended classes, you must state your assumption about the boundary
n = 454 customers
Frequency in thousands · Source: ONS 2014
EXERCISE 1E · Q1 · HISTOGRAMS & FREQUENCY DENSITY
Q1 — Supermarket Spending: Completed Table & Histogram
KEY FORMULA
FD = Frequency ÷ Class Width
Class (£P) | Freq | LCB | UCB | Width | FD |
P < 25 | 75 | 0 | 25 | 25 | 3.00 |
25 ≤ P < 50 | 97 | 25 | 50 | 25 | 3.88 |
50 ≤ P < 100 | 165 | 50 | 100 | 50 | 3.30 |
100 ≤ P < 150 | 86 | 100 | 150 | 50 | 1.72 |
150 ≤ P < 250 | 23 | 150 | 250 | 100 | 0.23 |
P ≥ 250 * | 8 | 250 | 350* | 100* | 0.08 |
Total | 454 | — | — | — | — |
* Assumption: P ≥ 250 class ends at £350 (width = 100). This must be stated clearly in your answer.
KEY POINTS
Gold values = answers to fill in. Teal = given in question.
FD = Frequency ÷ Width: e.g. 165 ÷ 50 = 3.30
Area of each bar = Frequency (not height). Bars have no gaps .
Highest FD bar is 50≤P<100 (FD=3.30) — most customers spend £50–£100.
HISTOGRAM — SUPERMARKET SPENDING (N = 454 CUSTOMERS)
EXERCISE 1E · Q2 · HISTOGRAMS & FREQUENCY DENSITY
Q2 — UK Population Age Distribution (June 2013)
KEY FORMULA
FD = Frequency (000s) ÷ Class Width
Age Group | Freq (000s) | Width | FD |
0 – 4 | 4,014 | 5 | 802.8 |
5 – 15 | 11,179 | 11 | 1016.3 |
16 – 44 | 21,453 | 29 | 739.8 |
45 – 64 | 16,328 | 20 | 816.4 |
65 – 74 | 6,031 | 10 | 603.1 |
75 – 89 | 4,574 | 15 | 304.9 |
90+ * | 527 | 10* | 52.7 |
Total | 64,106 | — | — |
* Assumption: 90+ group ends at age 100 (width = 10). State this clearly in your answer.
KEY POINTS
5–15 group : width = 16 − 5 = 11 → FD = 11,179 ÷ 11 = 1016.3
16–44 group : width = 45 − 16 = 29 → FD = 21,453 ÷ 29 = 739.8
Tallest bar: 5–15 (FD=1016.3) — large group due to 11-year span.
Unequal widths → must use FD , not frequency, for bar height.
HISTOGRAM — UK POPULATION BY AGE GROUP (SOURCE: ONS 2014)
Exercise 1E — Q3, Q4 & Q6
Frequency Density Tables & Histograms | AQA Chapter 1E
IQ Interval | Freq | Width | FD |
85 < x ≤ 95 | 21 | 10 | ? |
95 < x ≤ 100 | 76 | 5 | ? |
100 < x ≤ 105 | 137 | 5 | ? |
105 < x ≤ 110 | 129 | 5 | ? |
110 < x ≤ 115 | 108 | 5 | ? |
115 < x ≤ 125 | 83 | 10 | ? |
125 < x ≤ 135 | 16 | 10 | ? |
Total | 570 | — | — |
Q3 — IQ SCORES (N = 570)
Gainsby College student IQ distribution
TASKS
(a)(i) Calculate estimate for mean IQ and standard deviation.
(a)(ii) Last year: mean = 106, SD = 7.4. Write two comparison statements.
(b)(i) Draw a histogram using frequency density.
(b)(ii) Use histogram to estimate % with IQ > 106.
Delay (min) | Flights | Width | FD |
Early (−15 to 0) | 12,618 | 15 | ? |
1 – 15 late | — | 15 | ? |
16 – 30 late | 2,460 | 15 | ? |
31 – 60 late | 1,589 | 30 | ? |
61 – 180 late | 1,029 | 120 | ? |
181 – 360 late | 216 | 180 | ? |
> 360 late | 30 | — | ? |
Q4 — GATWICK FLIGHT DELAYS
Civil Aviation Authority data — Will's histogram is incorrect
WILL'S ERROR
Will used bar height = frequency instead of area = frequency . Class widths are unequal, so frequency density must be used.
TASKS
(a) Give reasons why Will's histogram is incorrect.
(b) Draw a correct histogram using frequency density.
Width (mm) | Doors | Width | FD |
Less than 750 | 0 | — | 0 |
750 – 790 | 24 | 45* | ? |
800 – 890 | 146 | 90 | ? |
900 – 990 | 124 | 90 | ? |
1000 – 1090 | 68 | 90 | ? |
1100 and over | 21 | — | ? |
Total | 383 | — | — |
Q6 — DOOR WIDTHS (N = 383)
Local authority doors, rounded to nearest 10 mm
BOUNDARY ASSUMPTION *
750–790 group: widths rounded to nearest 10 mm, so boundaries are 745 – 795 →width = 50 mm (not 40). State this assumption clearly.
TASKS
(a) Estimate number of doors not meeting official guideline (≥ 775 mm).
(b) Estimate % of doors the LA wants to widen (target: ≥ 850 mm).
Guidelines: Official ≥ 775 mm | LA target ≥ 850 mm
FD = Frequency ÷ Class Width
Area under histogram = Frequency
Unequal class widths → must use FD (not frequency) for histogram height
Always state assumptions for open-ended or ambiguous class boundaries
EXERCISE 1E · Q3 · IQ SCORES (N = 570)
Q3 — Gainsby College IQ Scores: FD Table, Histogram & Statistics
IQ Interval | Freq | Mid | Width | FD |
85 < x ≤ 95 | 21 | 90 | 10 | 2.1 |
95 < x ≤ 100 | 76 | 97.5 | 5 | 15.2 |
100 < x ≤ 105 | 137 | 102.5 | 5 | 27.4 |
105 < x ≤ 110 | 129 | 107.5 | 5 | 25.8 |
110 < x ≤ 115 | 108 | 112.5 | 5 | 21.6 |
115 < x ≤ 125 | 83 | 120 | 10 | 8.3 |
125 < x ≤ 135 | 16 | 130 | 10 | 1.6 |
Total | 570 | — | — | — |
MEAN IQ
103.4
Σfx ÷ 570 = 58,965 ÷ 570
STD DEV
≈ 8.0
Use σₙ₋₁ on calculator
COMPARISON WITH LAST YEAR (MEAN=106, SD=7.4)
This year's mean ( 103.4 ) is lower than last year's (106)— students scored lower on average.
This year's SD (≈8.0 ) is slightly higher than last year's (7.4) — scores are more spread out.
% with IQ > 106: Bars from 105–135 = 129+108+83+16 = 336. But 106 is 1/5 into the 105–110 bar, so approx 4/5 ×129 + 108 + 83 + 16 ≈ 310 ÷ 570 ≈ 54%
HISTOGRAM — IQ SCORE DISTRIBUTION (N = 570 STUDENTS) · VERTICAL LINE AT IQ = 106
EXERCISE 1E · Q4 · GATWICK FLIGHT DELAYS — WILL'S ERROR & CORRECT HISTOGRAM
Q4 — Gatwick Delays: Why Will Was Wrong & Correct FD Histogram
Delay (min) | Flights | Width | FD = f÷w |
Early (−15 to 0) | 12,618 | 15 | 841.2 |
1–15 late | not given | 15 | — |
16–30 late | 2,460 | 15 | 164.0 |
31–60 late | 1,589 | 30 | 53.0 |
61–180 late | 1,029 | 120 | 8.6 |
181–360 late | 216 | 180 | 1.2 |
>360 late | 30 | open* | excluded |
(A) WHY WILL'S HISTOGRAM WAS WRONG
Error 1: Will used frequency as bar height — but class widths are unequal, so this misrepresents the data.
Error 2: In a histogram, area = frequency , not height. Height must be frequency density (FD = f ÷ width).
Example: The 61–180 min bar (width=120) would appear far too tall if frequency (1,029) is used as height instead of FD (8.6).
* Note: The >360 min class has no upper boundary given — it cannot be plotted on a histogram without making an assumption. State this clearly and exclude it, or assume a boundary (e.g., 540 min).
CORRECT HISTOGRAM — GATWICK FLIGHT DELAYS (FD ON Y-AXIS, AREA = FREQUENCY)
EXERCISE 1E · Q6 · DOOR WIDTHS (N = 383)
Q6 — Door Widths: FD Table, Histogram & Guideline Estimates
Width (mm) | Doors | Boundaries | Width | FD |
Less than 750 | 0 | — | — | 0 |
750–790 * | 24 | 745–795 | 50 | 0.48 |
800–890 | 146 | 795–895 | 90 | 1.62 |
900–990 | 124 | 895–995 | 90 | 1.38 |
1000–1090 | 68 | 995–1095 | 90 | 0.76 |
1100 and over * | 21 | 1095–1185* | 90* | 0.23 |
Total | 383 | — | — | — |
* Key Assumption: 750–790 group: widths rounded to nearest 10 mm → boundaries are 745–795 (width = 50 mm, not 40). State this clearly. 1100+ assumed to end at 1185 (width = 90).
(A) DOORS < 775 MM (OFFICIAL GUIDELINE)
≈ 14 doors
745–795 band: 30/50 of 24 doors below 775
= 0.6 × 24 = 14.4 ≈ 14 doors
(B) % OF DOORS < 850 MM (LA TARGET)
≈ 29.5%
All 745–795 (24) + 55/90 of 795–895 (146)
= 24 + 89.2 = 113.2 → 113÷383 = 29.5%
HISTOGRAM — DOOR WIDTHS (N = 383) · DASHED LINES AT 775 MM AND 850 MM GUIDELINES
1F: Choosing Statistical Methods
SUMMARY
Key Principle: Always think carefully about whether your chosen methods are appropriate and help communicate your message clearly to the audience.
AVERAGES & SPREAD
Mean & Standard Deviation
Uses all data values; sensitive to extreme values and outliers.
Median & IQR
Not affected by extreme values; does not use all data in calculation.
Mode & Range
Easy to find; range is heavily affected by extreme values.
CHARTS & DIAGRAMS
Bar Charts, Pictograms & Pie Charts
Qualitative & ungrouped discrete data; pie shows proportions, not frequencies.
Stem-and-Leaf Diagrams
Quantitative data; organises raw values, enables mode, median and quartiles.
Box & Whisker Plots
Shows min, max, median and quartiles; does not display individual frequencies.
Cumulative Frequency Diagrams
Continuous & grouped discrete; plot at upper class boundary to find quartiles.
Histograms
Continuous & grouped discrete; area = frequency; use FD for unequal class widths.
Consolidation Exercise 1 — Q1, Q2 & Q3
CHAPTER 1 · SAMPLING & DATA
Q1
SAMPLING METHODS
(a)(i) Random sample:
Every member of the population has an equal chance of being selected.
(a)(ii) Representative:
A sample that reflects the characteristics of the whole population.
(b)(i) Cluster:
e.g. Select 3 schools at random, then survey all students in those schools.
(b)(ii) Quota:
e.g. Interview 10 males and 10 females aged 18–25 in a shopping centre.
Group | Count | Calculation | Sample |
Year 1 Male | 17 | 17/90 × 20 = 3.8 | 4 |
Year 1 Female | 12 | 12/90 × 20 = 2.7 | 3 |
Year 2 Male | 19 | 19/90 × 20 = 4.2 | 4 |
Year 2 Female | 11 | 11/90 × 20 = 2.4 | 2 |
Year 3 Male | 23 | 23/90 × 20 = 5.1 | 5 |
Year 3 Female | 8 | 8/90 × 20 = 1.8 | 2 |
Total | 90 | | 20 ✓ |
Q2
STRATIFIED SAMPLE — ENGINEERING COURSE (N = 90, SAMPLE = 20)
(b) Suggest two ways to select students from each group (e.g. random number table, lottery method).
Q3
PLANT HEIGHTS (CM) — OUTDOORS VS INDOORS
OUTDOORS (N = 22)
27, 32, 16, 10, 18, 13, 25, 23, 9, 11,
17, 21, 20, 17, 18, 29, 29, 15, 22, 11, 26, 31
INDOORS (N = 20)
17, 22, 23, 27, 35, 18, 27, 14, 32, 10,
27, 28, 30, 34, 12, 15, 22, 11, 26, 31
(a)(i) Draw a back-to-back stem-and-leaf diagram for both groups.
(a)(ii) Write down the mode for each group.
(a)(iii) Work out the range for each group.
(b)(i) Draw a box and whisker plot for each group.
(b)(ii) Describe and compare the two distributions.
Consolidation Exercise 1 — Q4 & Q5
Data Tables & Tasks
Key Reminder — Q5: Class widths are unequal. You must use Frequency Density = Frequency ÷ Class Width on the vertical axis of your histogram. Do not plot raw frequencies.
Q4 — French Test Marks
A Level French, n = 64 students
Mark Range | Frequency | Cumulative Freq. |
1 – 10 | 0 | 0 |
11 – 20 | 8 | 8 |
21 – 30 | 11 | 19 |
31 – 40 | 23 | 42 |
41 – 50 | 17 | 59 |
51 – 60 | 5 | 64 |
Total | 64 | — |
TASKS
a(i)
Draw a cumulative frequency graph for the data.
a(ii)
Find the 40th percentile. What does this value tell you?
a(iii)
The pass mark is 40%. Estimate the number of students who passed.
b
Draw a box and whisker plot to illustrate the distribution of marks.
Q5 — Clothes Spending
Julie's survey, n = 88 students
Amount Spent (£P) | Frequency | Class Width |
P < 20 | 8 | 20 |
20 ≤ P < 30 | 19 | 10 |
30 ≤ P < 40 | 27 | 10 |
40 ≤ P < 60 | 18 | 20 |
60 ≤ P < 100 | 14 | 40 |
P ≥ 100 | 2 | — |
Total | 88 | — |
TASKS
a
Julie says "My data is continuous, secondary data." Is she correct? Explain your answer fully.
b
Draw a histogram for Julie's data. Remember: unequal class widths require frequency density.
c
Find an estimate for the mean amount spent on clothes and the standard deviation.
Consolidation Exercise 1 — Q6, Q7 & Q8
AQA MATHEMATICAL STUDIES
Q6 — PETROL & DIESEL CONSUMPTION GRAPHS
CONTEXT
UK consumption of petrol and diesel, 2000–2013. Two graphs are provided:
TANYA'S GRAPH
No scale on y-axis; uses misleading pictogram-style bars — bar widths vary, making comparison unreliable.
OLIVER'S GRAPH
Proper numerical scale on y-axis; consistent bar widths; clearly labelled axes.
TASKS
(a) List the ways in which each diagram can be improved.
(b) Describe how the consumption of petrol and diesel has changed over this time period.
Consider: missing title, missing units, no key, misleading scale, pictogram distortion.
Q7 — BLACKPOOL SCHOOLS SAMPLING
BLACKPOOL LOCAL AUTHORITY
Total: 63 schools and colleges
39
Primary
Schools
17
Secondary
Schools
7
Colleges
(16–18)
PROPOSED METHOD
Choose one school/college from each group at random, then interview two teachers from each selected institution.
TASKS
(a) Give two reasons why this is not a good sample.
(b) Describe a better sampling method for this situation.
Hint: consider sample size, proportional representation, and whether two teachers per school is sufficient.
Q8 — BOX OFFICE TAKINGS (BFI 2014) — HISTOGRAM
Takings T (£m) | Frequency | Class Width |
0 < T < 0.1 | 433 | 0.1 |
0.1 ≤ T < 1 | 133 | 0.9 |
1 ≤ T < 5 | 70 | 4 |
5 ≤ T < 10 | 28 | 5 |
10 ≤ T < 20 | 20 | 10 |
20 ≤ T < 30 | 6 | 10 |
30 ≤ T < 50 | 8 | 20 |
Total | 698 | — |
DATA TABLE — 698 FILMS
Unequal class widths — use
frequency density
for histogram (fd = freq
÷ class width)
TASKS
(a) Write down the modal class .
(b) Estimate the mean and standard deviation .
(c) Estimate the median and IQR . Describe any problems.
(d) Write three sentences explaining what the data shows.
Source: BFI Statistical Yearbook 2014
Consolidation Exercise 1 — Q9 & Q10
AQA Mathematical Studies
Q9 — HOSPITAL WAITING TIMES
Waiting Time | % of Patients |
Less than 5 weeks | 2% |
5–9 weeks | 17% |
10–15 weeks | 26% |
16–17 weeks | 38% |
18 weeks (target boundary) | 12% |
19 weeks | 4% |
20 weeks | 1% |
More than 20 weeks | 0% |
Non-emergency treatment waiting times (Target: no more than 18 weeks)
Highlighted row = 18-week target boundary
TASKS
▸ Comment on the hospital's performance against the 18-week target.
▸ Use statistical measures and/or diagrams to support your comments.
Hint: Consider cumulative frequencies, median, and what % waited ≤ 18 weeks.
Q10 — UK GENDER GAP: YEARS OF LIFE LOST
Age Group | F Yrs Lost | M Yrs Lost | F Pop | M Pop | F Deaths | M Deaths |
20–24 | 5,544.0 | 4,582.0 | 1,774,400 | 1,829,400 | 90 | 79 |
30–34 | 12,873.3 | 13,620.6 | 1,850,300 | 1,831,800 | 249 | 282 |
50–54 | 55,094.0 | 65,072.7 | 1,824,700 | 1,792,800 | 1,690 | 2,191 |
70–74 | 89,900.0 | 111,577.5 | 1,106,900 | 998,900 | 5,800 | 8,265 |
Premature death data by age group — UK Gender Gap Report (Oct 2014). UK ranked 26th (WEF).
TASKS
(a) How does the mean number of years lost by females in each age group compare with males?
(b) Write down two more questions you could ask based on this data.
Hint: Calculate mean years lost per death for each gender and age group.
Source: Health & Social Care Information Centre (HSCIC)
Q11 — Gatwick vs Heathrow: CAA Passenger Survey 2013
Consolidation Exercise 1
YOUR TASK
Use the data in all three tables to compare the use of Gatwick and Heathrow by passengers who travel by air. Consider group sizes, trip lengths, and age profiles. Use statistical measures (mean, median, percentages) and diagrams to support your comparisons.
Extension: Use the internet to find and compare data for other UK airports (e.g. Manchester, Birmingham, Edinburgh).
SOURCE
Civil Aviation Authority (CAA) Passenger Survey Report 2013.
Gatwick: 32,402,000 passengers | Heathrow: 45,744,000 passengers
No. in Group | Gatwick % | Heathrow % |
1 | 36.8 | 63.3 |
2 | 36.1 | 23.9 |
3 | 4.9 | 4.5 |
4 | 12.8 | 4.1 |
5 | 3.7 | 1.2 |
6+ | 5.8 | 3.1 |
Total passengers (000s) | 32,402 | 45,744 |
TABLE 1 — TRAVELLING GROUP SIZE
Length (x days) | Gatwick % | Heathrow % |
x < 1 | 2.7 | 4.7 |
1 ≤ x < 3 | 11.4 | 13.1 |
3 ≤ x < 6 | 30.4 | 25.1 |
6 ≤ x < 8 | 24.3 | 12.7 |
8 ≤ x < 15 | 23.4 | 21.8 |
x ≥ 15 | 7.8 | 22.7 |
Total | 100% | 100% |
TABLE 2 — TRIP LENGTH (DAYS)
Age Group | Gatwick % | Heathrow % |
10 or under | 3.2 | 0.9 |
11–15 | 3.4 | 1.6 |
16–19 | 4.9 | 4.2 |
20–24 | 10.5 | 9.8 |
25–34 | 18.5 | 24.0 |
35–44 | 16.0 | 20.3 |
45–54 | 17.9 | 18.5 |
55–59 | 7.6 | 7.3 |
60–64 | 8.3 | 6.1 |
65–74 | 8.0 | 5.9 |
Over 74 | 1.7 | 1.4 |
Total | 100% | 100% |
TABLE 3 — AGE DISTRIBUTION
Exercise 1F — Activity 17 & 18
RAINFALL DATA
England & Wales — Hadley Centre Data
Year | Rainfall (mm) | Note |
1770 | 1079.4 | Highest |
1771 | 792.9 | Lowest |
1772 | 1031.8 | |
1773 | 1033.8 | |
1774 | 994.7 | |
1775 | 1012.0 | |
1776 | 847.4 | |
1777 | 860.3 | |
1778 | 887.9 | |
1779 | 900.5 | |
TABLE 1 — 1770 TO 1779 (MM)
Year | Rainfall (mm) | Note |
2000 | 1232.5 | Highest |
2001 | 970.0 | |
2002 | 1117.8 | |
2003 | 761.4 | Lowest |
2004 | 973.6 | |
2005 | 825.1 | |
2006 | 904.8 | |
2007 | 1022.7 | |
2008 | 1089.6 | |
2009 | 977.1 | |
TABLE 2 — 2000 TO 2009 (MM)
ACTIVITY 18 — STUDENTS' METHODS
Ahmed
Bar chart comparing monthly rainfall for two years a century apart (1766 vs 2014)
Kayleigh
Time series graph of annual rainfall 1766–2010 (long-term trend)
Cilla
Mean & SD of monthly rainfall for selected years
Kirsty
Pie charts for first and last years (1766 & 2014)
Tim
Median & IQR for 30-year groups (1771–2010)
FOR EACH STUDENT, EVALUATE:
(a)
Was the method appropriate?
(b)
Describe any patterns or changes found
(c)
How could their work be improved?
(d)
What other methods could they have used?
ANNUAL RAINFALL COMPARISON — 1770S VS 2000S (MM)
Activity 18 — Students' Statistical Methods
EXERCISE 1F
FIVE STUDENTS INVESTIGATED HADLEY CENTRE RAINFALL DATA (1766 ONWARDS):
Ahmed
Bar chart — two years compared
Kayleigh
Time series — 1766–2010
Cilla
Mean & SD per year
Kirsty
Pie charts — 1766 & 2014
EVALUATE EACH STUDENT'S WORK:
(a) Was the method appropriate for investigating rainfall changes over time?
(b) Describe any patterns or changes found in the data.
(c) How could their work be improved?
(d) What other statistical methods could they have used?
CILLA — MEAN & SD OF MONTHLY RAINFALL (MM)
Year | Mean (mm) | SD (mm) |
1766 | 63.0 | 33.8 |
1767 | 78.4 | 32.0 |
1768 | 103.9 | 38.7 |
1769 | 79.2 | 27.9 |
1770 | 90.0 | 41.0 |
2010 | 68.5 | 25.1 |
2011 | 65.6 | 28.5 |
2012 | 103.7 | 46.9 |
2013 | 76.4 | 35.4 |
2014 | 92.1 | 45.2 |
TIM — MEDIAN & IQR OF ANNUAL RAINFALL BY 30-YEAR GROUP
Period | Median (mm) | IQR (mm) |
1771–1800 | 888.35 | 199.6 |
1801–1830 | 903.55 | 157.35 |
1831–1860 | 873.95 | 169.3 |
1861–1890 | 912.45 | 183.95 |
1891–1920 | 907.45 | 155.45 |
1921–1950 | 925.4 | 165.65 |
1951–1980 | 904.5 | 206.65 |
1981–2010 | 971.8 | 170.25 |
Chapter Summary — Key Points & Review
CHAPTER 1
Data Types
Know the meaning of key terms for classifying data correctly before analysis.
Qualitative
Quantitative
Discrete
Continuous
Primary
Secondary
Sampling Methods
Appreciate strengths and limitations of each method; larger samples improve accuracy but cost more time and money.
Random
Cluster
Stratified
Quota
Stratified: (group size ÷ total) × sample size
Numerical Measures
Represent data numerically from raw data, frequency tables, or statistical diagrams.
Mean
Median
Mode
Quartiles
IQR
Std Dev
Mean from grouped data: use midpoints of each class. Use σₙ₋₁ on calculator.
Statistical Diagrams
Represent data appropriately and interpret diagrams to reach valid conclusions.
Histograms
CF Graphs
Stem-and-Leaf
Box Plots
Key Reminders
Critical rules to avoid common errors in exams and coursework.
CF graphs: always plot at upper class boundary
Histograms: FD = frequency ÷ class width ; area = frequency
Percentiles: nth percentile at n% × total
Essential Formulae
Core formulae required for all Chapter 1 calculations and exam questions.
Stratified sample: (group ÷ total) × n
nth percentile: n% × Σf on CF graph
Grouped mean: Σ(midpoint × f) ÷ Σf
AQA MATHEMATICAL STUDIES — CHAPTER 1 COMPLETE
Chapter 1 Complete
Analysis of Data — from data types and sampling through to histograms and consolidation
TOPICS MASTERED
Cumulative Frequency
Graphs & percentiles
Histograms
Frequency density
Statistical Methods
Choosing appropriately
Consolidation Exercise
All Chapter 1 topics
NEXT STEPS
Practise past AQA exam questions on data analysis
Use your calculator for mean and SD from grouped data
Review the key formulae card before your next lesson
WELL
DONE!
1E: Histograms — Frequency Density
AQA MATHEMATICAL STUDIES
Speed (mph) | Freq (f) | Width (w) | FD = f ÷ w |
0 < x ≤ 10 | 1 | 10 | 0.10 |
10 < x ≤ 20 | 8 | 10 | 0.80 |
20 < x ≤ 26 | 6 | 6 | 1.00 |
26 < x ≤ 30 | 28 | 4 | 7.00 |
30 < x ≤ 32 | 14 | 2 | 7.00 |
32 < x ≤ 34 | 3 | 2 | 1.50 |
34 < x ≤ 36 | 0 | 2 | 0.00 |
36 < x ≤ 38 | 2 | 2 | 1.00 |
38 < x ≤ 40 | 1 | 2 | 0.50 |
SCHOOL TRAFFIC SURVEY — SPEED DATA (N = 65 VEHICLES)
Highlighted rows = peak FD (speed limit zone 26–32 mph)
HISTOGRAM — FREQUENCY DENSITY VS SPEED
KEY FORMULA
FD = Frequency ÷ Class Width
To read back: Frequency = FD × Class Width
GOLDEN RULE
AREA represents frequency — not height
Use for continuous data with unequal class widths. No gaps between bars.
COMMON MISTAKE
Tom's diagram was misleading
He used HEIGHT not AREA to represent frequency — this is incorrect for unequal widths.
CHAPTER 1
Summary — Key Points & Review
Key Reminders:
Always compare average AND spread
Always check for outliers using fences
Plot CF at upper class boundary
Use FD for unequal class width histograms
Data Types
Qualitative — descriptive categories.
Discrete — exact countable values.
Continuous — any value in a range.
Sampling
Random, cluster, quota methods.
Stratified sample size per group:
(group ÷ total) × sample size
Averages
Mode — most frequent value.
Median — middle value when ordered.
Mean = Σx/n or Σfx/Σf
Spread & Outliers
Use σₙ₋₁ on calculator for SD.
IQR = UQ − LQ
Outlier: < LQ−1.5×IQR or > UQ+1.5×IQR
Box Plots
Five-number summary: Min, LQ, Median, UQ, Max.
Mark outliers with ×. Always compare median (average) and IQR (spread) when comparing two distributions.
Cumulative Frequency
Plot CF against upper class boundary . Draw smooth S-curve.
Median = n/2 · LQ = n/4 · UQ = 3n/4
Histograms
Used for continuous data with unequal class widths .
Area represents frequency — not height.
FD = Frequency ÷ Class width