1 of 41

CHAPTER 1

Analysis of Data

AQA Mathematical Studies  ·  Paper 1

1.1 Data Types & Sampling

1.2 Averages & Stem-and-Leaf

1.3 Measures of Spread

1.4 Box & Whisker Plots

1.5 Cumulative Frequency

1.6 Histograms

AQA LEVEL 3 CERTIFICATE

Mathematical Studies · Paper 1

2 of 41

AQA MATHEMATICAL STUDIES — CHAPTER 1

Chapter Overview & Learning Objectives

SECTION 1.1

Data Types & Sampling

Qualitative, discrete & continuous data; random, stratified, cluster and quota sampling methods.

SECTION 1.2

Averages & Stem-and-Leaf

Mean, mode and median from tables; back-to-back stem-and-leaf diagrams for comparison.

SECTION 1.3

Measures of Spread

Range, interquartile range and standard deviation; identifying outliers using the 1.5 × IQR rule.

SECTION 1.4

Box & Whisker Plots

Five-number summary (min, Q1, median, Q3, max); plotting outliers; comparing distributions.

SECTION 1.5

Cumulative Frequency

Drawing CF graphs; reading percentiles; estimating median and IQR directly from the graph.

SECTION 1.6

Histograms

Frequency density on the y-axis; handling unequal class widths; interpreting histogram area.

Prerequisites: Frequency tables, pie charts, bar charts, and percentage calculations are assumed knowledge for this chapter.

3 of 41

1.1 Data Types & Sampling Methods

AQA Mathematical Studies — Chapter 1 · Paper 1 Topic

Data Types

THREE CATEGORIES

Qualitative (Categorical)

Descriptive — not a number

colour

gender

exam grade

nationality

Quantitative — Discrete

Exact countable values only

no. of texts

shoe size

ticket prices

age

Quantitative — Continuous

Any value within a range

height

weight

time

leaf area

Quick test: "Can it take any value in a range?" Continuous. "Only exact/countable values?"Discrete. "A description?" Qualitative.

Sampling Methods

Population — all members of the group being studied

Sample — sub-group used when population too large or costly

FOUR METHODS

Random

Every member has an equal chance of selection — unbiased but can be expensive

Stratified

Groups sampled in same proportion as population — most representative

Stratified Formula:

Sample from group = (group size ÷ total) × sample size

Cluster

Randomly select natural sub-groups (e.g. Local Authorities), then survey all within

Quota

Interviewers fill set quotas of types — quick but may introduce bias

Worked Example — Stratified Sample of 20 from 1522 students

Female Part time (465):

465÷1522×20 =

6

Female FullT (236):

236÷1522×20 =

3

Male Part Time (624):

624÷1522×20 =

8

Male FullT (197):

197÷1522×20 =

3

Check: 6 + 3 + 8 + 3 = 20 (must equal total sample size)

4 of 41

Stratified Sampling — Worked Examples

FORMULA

Sample from group = (Group size ÷ Total population) × Sample size

Round to nearest whole number · Check totals sum correctly

Key Rule: Always round each group to the nearest whole number, then check all group samples add up to the total sample size before finalising.

1

College Sample

Total = 1522 · n = 20

1522 students

Sample size: 20

4 groups

Female Part-Time

465 ÷ 1522 × 20 =

6

Female Full-Time

236 ÷ 1522 × 20 =

3

Male Part-Time

624 ÷ 1522 × 20 =

8

Male Full-Time

197 ÷ 1522 × 20 =

3

Verification: 6 + 3 + 8 + 3 = 20

Matches sample size

2

Company Sample

Total = 922 · n = 200

Department

M FT

M PT

F FT

F PT

Total

Factory

239

171

179

91

680

Warehouse

55

20

28

15

118

Office

29

7

42

5

83

Delivery

17

4

1

20

42

922 employees

Sample size: 200

4 departments

DEPARTMENT BREAKDOWN

WORKED CALCULATION — FACTORY MALE FT

239 ÷ 922 × 200 = 52

Apply the same formula to every group. All results must sum to 200.

All group samples must sum to 200

Always verify totals

5 of 41

Exercise 1A — Skills

SKILL GUIDE

DATA TYPE CLASSIFICATION

Qualitative

Descriptive — not a number

colour, gender, nationality, grade

Discrete

Countable exact values only

number of texts, shoe size, ticket prices

Continuous

Any value in a range (measured)

weight, height, time, length, area

STRATIFIED SAMPLING — METHOD

1

Find the total population size

2

Apply the formula for each group:

Sample from group = (Group size ÷ Total) × n

3

Round each result to the nearest whole number

4

Check: all group samples must sum to total sample size

Group

Group Size

Total

Sample Size (n)

Calculation

Result

Year 12

90

240

40

(90 ÷ 240) × 40

= 15

Year 13

150

240

40

(150 ÷ 240) × 40

= 25

WORKED EXAMPLE — STRATIFIED SAMPLING

Example: A school has 240 students. Year 12 = 90, Year 13 = 150. Take a stratified sample of 40.

  Verification: 15 + 25 = 40  (totals match sample size)

SURVEY METHODS — KEY POINTS

Random

Every member has equal chance of selection. Unbiased but can be expensive and time-consuming.

Stratified

Groups sampled in proportion to population. Representative of all subgroups. Use the formula.

Cluster

Randomly select natural sub-groups (e.g. schools, regions), then survey all within chosen clusters.

Quota

Interviewers fill set quotas by type. Quick and practical but may introduce selection bias.

Exam Tip: Always show your calculation clearly and verify the total. If rounding causes the sum to be off by 1, adjust the largest group. State which method you used and why it is appropriate.

6 of 41

Exercise 1A — Questions

ATTEMPT ALL QUESTIONS

Work through all six questions before checking answers on the next slide.

Q1

DATA TYPES

Number of texts sent

Colour of flowers

Weight of puppies

Exam grade

Shoe size

Length of feet

Time asleep

Leaf area

Ticket prices

Gender of fans

Age of fans

Nationality of players

Classify each item as Qualitative , Discrete or Continuous :

Q2

DATA SOURCES

(a) Give one example of a primary data source you could use.

(b) Give one example of a secondary data source you could use.

Describe how you would find both primary and secondary data on students' earnings from part-time work.

Q3

SURVEY METHODS

(a) Suggest two questions the company might ask.

(b) Comment on three methods: bus station interviews , postal questionnaires , and radio/TV invitation to respond .

A market research company surveys people about a new cinema complex .

Q4

STRATIFIED SAMPLING — COMPANY SURVEY

Formula: (Group size ÷ 922) × 200

(a) Calculate the sample size for Factory Male FT . Show your working.

(b) Calculate the sample sizes for all other groups .

(c) Verify your answers sum to 200.

A company wants a stratified sample of 200 employees from its 922 staff . The two-way table shows the breakdown by department and employment type.

Department

Male FT

Male PT

Female FT

Female PT

Total

Factory

239

171

179

91

680

Warehouse

55

20

28

15

118

Office

29

7

42

5

83

Delivery

17

4

1

20

42

TOTAL

340

202

250

131

922

Q5

QUOTA SAMPLING

(a) State what quotas you would set and why.

(b) Explain how interviewers would fill each quota.

(c) Give one advantage and one disadvantage of this method.

Describe how quota sampling could be used to survey 500 people about a proposed new airport .

Q6

CLUSTER SAMPLING

Describe how cluster sampling could be used to investigate health care provision across England .

7 of 41

Exercise 1A — Answers

AQA CHAPTER 1.1

Q1

DATA TYPE CLASSIFICATION

Qualitative

Discrete

Continuous

Colour of flowers

No. of texts sent

Weight of puppies

Exam grade

Shoe size

Length of feet

Gender of fans

Ticket prices

Time asleep

Nationality of players

Age of football fans

Leaf area

Qualitative = descriptions  |  Discrete = countable exact values  |  Continuous = any value in a range

Q4

STRATIFIED SAMPLING — 1200 STUDENTS, SAMPLE OF 50

Formula: (Group Size ÷ Total) × Sample Size

Year 12 (480):

(480 ÷ 1200) × 50 =

20 students

Year 13 (720):

(720 ÷ 1200) × 50 =

30 students

Verification: 20 + 30 = 50

Selection: use random numbers or names in a hat within each year group

Q2

DATA SOURCES — STUDENT EARNINGS

PRIMARY DATA

Conduct your own survey or questionnaire asking students directly about their part-time earnings.

SECONDARY DATA

Use existing records from HMRC, college databases, or government statistics on student employment.

Q3

SURVEY METHOD CRITIQUE — CINEMA COMPLEX

Bus Station

Biased — only reaches evening/Saturday users; not representative of all potential cinema-goers.

Postal

Low response rate — may not reach all areas or demographics equally; self-selection by motivated respondents.

Radio/TV

Self-selection bias — only motivated or opinionated people respond; not representative of general public.

Q5

QUOTA SAMPLING   

Q6

CLUSTER SAMPLING

Q5 — QUOTA SAMPLING (AIRPORT)

Set quotas by age and gender in proportion to the local population. Interviewers select participants until each quota is filled.

Not random — interviewers choose who to approach.

Q6 — CLUSTER SAMPLING (HEALTH CARE)

Randomly select a sample of Local Authorities (clusters), then survey all people within those LAs about health care provision.

Not a random sample of England's entire population.

8 of 41

1.2 Averages: Mean, Mode & Median

AQA CHAPTER 1.2

WORKED EXAMPLE — DRIVING TEST AGES

GROUP

MODE

MEAN

MEDIAN

NOTE

Female

17

19

Outlier age 49 inflates mean

Male

18

19.9

19.5

All three averages close together

Key Insight: The female mean (22.7) is distorted by the outlier age 49 — the median (19) better represents the typical female candidate. When outliers are present, the median is more reliable than the mean.

Mode & Median

Mode — easy to find; always a data value; may not exist or may be bimodal

Mode — does not use all data values in its calculation

Median — often a data value; not affected by extreme values (outliers)

Median — does not use all data; robust to skewed distributions

Mean

x̄ = Σx ÷ n

Uses all data values — every value contributes to the result

Affected by extreme values — outliers pull the mean up or down significantly

Involves more calculation — sum all values then divide by n

Not usually a data value — result is often a decimal

22.7

distorted

9 of 41

Example 1 & 2: Stem-and-Leaf Diagrams

AQA Chapter 1.2 — Constructing and interpreting stem-and-leaf diagrams

Example 1 — Test Marks (out of 60)

DATA: 39, 37, 56, 44, 32, 40, 26, 58, 31, 42, 37, 51, 29, 38, 28, 42  (N = 16)

STEM | LEAF  (ORDERED)

2

6  8  9

3

1  2  7  7  8  9

4

0  2  2  4

5

1  6  8

Key:  5 | 1  means 51 marks  ·  n = 16

Mode

37

 and 

42

 — bimodal (each appears twice)

Median

= (8th + 9th) ÷ 2 = (38 + 39) ÷ 2 =

38.5 marks

n

16 values  →  middle pair: 8th & 9th

Example 2 — Driving Test Ages (Back-to-Back)

Female (leaves)

Stem

(leaves) Male

9  9  8  7  7  7

1

7  7  8  8  8  8  9

6  5  2  1

2

0  0  1  2  2  3  6

9

4

 —

Key:  5 | 2 | 6  means Female 25 , Male 26

Female

Median =

19

 ·  Range =

32

(outlier: age 49)

Male

Median =

19.5

 ·  Range =

9

Female ages are far more widely spread than male ages — the outlier age 49 inflates

the female range to 32, compared to just 9 for males. Both groups share a similar median

(~19).

Always include a Key

— it must be shown on every stem-and-leaf diagram.

10 of 41

AQA MATHEMATICAL STUDIES — CHAPTER 1.3

Measures of Spread: Range, IQR & Standard Deviation

Range = Max − Min

IQR = UQ − LQ

Use σₙ₋₁ on calculator

Key Insight: Both classes share the same median (59) , but Class B is far more spread out— IQR = 33 vs IQR = 7 for Class A. Same centre, very different spread.

Standard Deviation: Use the σₙ₋₁ button on your calculator (or STDEV in Excel). A higher σmeans the data is more spread out from the mean.

Measure

Class A (n=15)

Class B (n=14)

Lower Quartile (LQ)

56

37

Median

59

59

Upper Quartile (UQ)

63

70

IQR = UQ − LQ

7

33

Spread

Compact

Much more spread

CLASS A VS CLASS B — COMPARISON (N=15 AND N=14)

OUTLIER RULE & WORKED EXAMPLE

OUTLIER DEFINITION

Value < LQ − 1.5 × IQR

Value > UQ + 1.5 × IQR

CLASS A — OUTLIER CHECK

LQ − 1.5 × IQR = 56 − 1.5 × 7

= 56 − 10.5 = 45.5

Value 8 < 45.5

Value 8 IS an outlier

RANGE

Easy to find. Affected by extreme values (outliers). Uses only two data points.

IQR (INTERQUARTILE RANGE)

Uses middle 50% of data. Not affected by outliers. For n 20: split data either side of median.

STANDARD DEVIATION Σ

Uses ALL data. Gives fair comparison. More complex to calculate. Use STDEV in spreadsheet.

11 of 41

IQR Worked Examples 1 & 2

AQA CH 1.3

Key Insight: IQR is more robust than range — it ignores extreme values. Use the rule: outlier if value < LQ − 1.5×IQR or value > UQ + 1.5×IQR.

EXAMPLE 1

Class A Test Scores

n = 15

SORTED DATA

8 , 54, 55, 56, 58, 59, 59, 59 , 62, 62, 63, 63, 6 3, 64, 65

Median

8th value = 59

Lower half

8, 54, 55, 56, 58, 59, 59

LQ (4th)

= 56

Upper half

62, 62, 63, 63, 63, 64, 65

UQ (4th)

= 63

IQR

63 − 56

= 7

OUTLIER CHECK

LQ − 1.5 × IQR = 56 − 10.5 = 45.5

Value 8 < 45.5

8 IS an outlier

EXAMPLE 2A

Group P

n = 11

SORTED DATA

18, 19, 19, 21, 23, 23 , 26, 27, 32, 35, 43

Median

6th value = 23

Lower half

18, 19, 19, 21, 23

LQ (3rd)

= 19

Upper half

26, 27, 32, 35, 43

UQ (3rd)

= 32

IQR

32 − 19

= 13

OUTLIER CHECK

LQ − 1.5 × IQR = 19 − 19.5 = −0.5

UQ + 1.5 × IQR = 32 + 19.5 = 51.5

All values within bounds

No outliers

EXAMPLE 2B

Group Q

n = 10

SORTED DATA

20, 21, 24, 25, 29, 31, 33, 35, 37, 69

Median

(29 + 31) ÷ 2 = 30

Lower half

20, 21, 24, 25, 29

LQ (3rd)

= 24

Upper half

31, 33, 35, 37, 69

UQ (3rd)

= 35

IQR

35 − 24

= 11

OUTLIER CHECK

UQ + 1.5 × IQR = 35 + 16.5 = 51.5

Value 69 > 51.5

69 IS an outlier

12 of 41

Exercise 1B — Skills

HOW TO FIND MEAN, IQR & STANDARD DEVIATION

FINDING THE MEAN

1

Add all values together: Σx

2

Divide by the number of values:

x̄ = Σx / n

Frequency table: multiply each value by its frequency, sum, then divide by total frequency

x̄ = Σfx / Σf

Weighted mean (combining groups):

(n₁x̄₁ + n₂x̄₂) / (n₁ + n₂)

Do NOT simply average the two means!

FINDING THE IQR

1

Sort data in ascending order

2

Find the median (middle value)

3

For n 20 : split data either side of median (exclude median itself)

4

LQ = median of lower half  |  UQ = median of upper half

5

IQR = UQ − LQ

Outlier rule:

value < LQ − 1.5×IQR  or  value > UQ + 1.5×IQR

STANDARD DEVIATION

Use the

σₙ₋₁

button on your calculator

In a spreadsheet: use the STDEV function

Higher σ = data is more spread out from the mean

Standard deviation uses ALL data values — gives a fair comparison between datasets

Use σ to compare spread when two datasets have the same mean but different variability

TARGET MEAN — WORKED EXAMPLE

Formula: Runs needed = (target mean × new n) − current total

Runs needed = (target x̄ × new n) − (current x̄ × old n)

JACK'S CRICKET SCORES

Jack's mean from 8 matches = 45

Target mean for 9 matches = 50

Current total = 8 × 45 = 360

Target total = 9 × 50 = 450

Runs needed in match 9:

450 − 360 = 90 runs

Always find both totals first, then subtract to find the missing value

13 of 41

Exercise 1B — Questions

ATTEMPT BEFORE CHECKING ANSWERS

Q1 — AVERAGES & RANGE

Find the mode, median, mean and range for the 20 Premiership season ticket prices below.

£1014

£335

£499

£595

£550

£544

£501

£365

£710

£299

£532

£383

£499

£608

£459

£400

£449

£795

£349

£640

Group

Number

Wage/hr

Packers

42

£9.50

Fork-lift drivers

14

£13.50

Q3 — WEIGHTED MEAN

Stacy claims the overall mean wage is £11.50/hr. Explain why she is wrong and calculate the correct mean.

Hint: Use weighted mean — do NOT simply average the two wages.

Q4 — TARGET MEAN

How many runs must Jack score in match 9 to achieve an overall mean of 50?

8

Matches played

45

Current mean

50

Target mean (9 matches)

?

Runs in match 9

Hint: Target total = target mean × new n. Runs needed = target total − current total.

Q5 — MEDIAN & IQR

For each dataset find the median and interquartile range (IQR).

(a) Bus late arrival times (minutes) — n = 15

0, 0, 0, 5, 5, 10, 10, 10, 15, 15, 20, 25, 30, 35, 40

(b) Seeds germinating — n = 15

3, 4, 4, 5, 5, 5, 6, 6, 7, 8, 9, 10, 11, 12, 15

(c) Daily lunch spending — n = 11

£2.50, £3.00, £3.50, £3.50, £4.00, £4.50, £5.00, £5.50, £6.00, £7.00, £8.00

Q6 — MEAN & STANDARD DEVIATION

Find the mean and standard deviation for each dataset. Use σn−1 on your calculator.

(a) Test marks — n = 10

35, 38, 40, 40, 42, 43, 44, 45, 46, 47

(b) Daily temperatures (°C) — n = 10

12, 14, 15, 16, 17, 18, 19, 20, 21, 23

Calculator tip: Enter data STAT mode use σ

n−1

(not σ

n

) for sample standard deviation.

Answers on the next slide

14 of 41

Exercise 1B — Answers

AQA CHAPTER 1.3

Weighted mean formula:

(n₁x̄₁ + n₂x̄₂) ÷ (n₁+n₂)

Target mean:

runs needed = (target mean × new n) − current total

Standard deviation: use

σₙ₋₁

button on calculator

Q1 / AVERAGES & RANGE

Season Ticket Prices (n = 20)

MODE

£499

MEDIAN

£500

MEAN

£526.30

RANGE

£715

1

Sorted data: £299 … £1014. Mode = £499 (appears twice)

2

Median = (10th + 11th) ÷ 2 = (£499 + £501) ÷ 2 = £500

3

Mean = £10,526 ÷ 20 = £526.30  |  Range = £1014 − £299 = £715

Q3 / WEIGHTED MEAN

Stacy's Error — 42 Packers & 14 Fork-lift Drivers

STACY'S (WRONG)

£11.50/hr

CORRECT MEAN

£10.50/hr

Stacy averaged the two means without weighting by group size

(42 × £9.50 + 14 × £13.50) ÷ 56 = (£399 + £189) ÷ 56 = £10.50/hr

Q4 / TARGET MEAN

Jack's Cricket Scores — Match 9 Target

RUNS NEEDED

90 runs

1

Current total = 8 × 45 = 360

2

Target total = 9 × 50 = 450

3

Runs needed = 450 − 360 = 90 runs

Q5(A) / MEDIAN & IQR

Bus Late Arrival Times (n = 15, minutes)

MEDIAN

10

LQ

5

UQ

25

IQR

20

1

Sorted: 0,0,0,5,5, 10 ,10,10,15,15,20,25,30,35,40. Median (8th) = 10

2

Lower half: 0,0,0,5,5,10,10 LQ (4th) = 5  |  Upper half UQ (12th) = 25

3

IQR = 25 − 5 = 20

Q6(A) / MEAN & STANDARD DEVIATION

Test Marks: 35, 38, 40, 40, 42, 43, 44, 45, 46, 47

MEAN (X̄)

42

STD DEV (Σ)

3.3

1

Σx = 35+38+40+40+42+43+44+45+46+47 = 420

2

Mean = 420 ÷ 10 = 42  |  Use σₙ₋₁ on calculator σ ≈ 3.3

15 of 41

1.4 Box and Whisker Plots

EXAMPLE 3

Manchester United Goals — 2013–14 Season (n = 38 matches)

AQA Mathematical Studies — Chapter 1.4 Example 3 | Man Utd 2013–14 Season

BOX PLOT DIAGRAM

Whisker drawn to 6 (last non-outlier). Value 7 marked with × as outlier.

0

Min

0.5

LQ

2

Median

3

UQ

6

Max*

0

1

2

3

4

5

6

7

×

Min=0

LQ=0.5

Med=2

UQ=3

Max=6

Outlier

Goals per Match

FREQUENCY TABLE — GOALS PER MATCH

= (9×0 + 9×1 + 9×2 + 7×3 + 4×4 + 0×5 + 2×6 + 1×7) ÷ 38

= (0 + 9 + 18 + 21 + 16 + 0 + 12 + 7) ÷ 38 = 64 ÷ 38

x̄ = 1.68 goals

Goals (x)

Freq (f)

Cum. Freq

fx

0

9

9

0

1

9

18

9

2

9

27

18

3

7

34

21

4

4

38

16

5

0

38

0

6

2

40

12

7 ×

1

41 — wait

7

Total

38

64

KEY STATISTICS

📊

MEDIAN (19TH VALUE)

2 goals

📦

LQ (9.5TH VALUE)

0.5 goals

📦

UQ (28.5TH VALUE)

3 goals

📐

MEAN X̄ = ΣFX/ΣF

1.68 goals

σ

STANDARD DEVIATION Σ

1.32 goals

IQR CALCULATION

IQR = UQ − LQ

= 3 − 0.5 = 2.5 goals

OUTLIER CHECK

Upper fence = UQ + 1.5 × IQR

= 3 + 1.5 × 2.5

= 3 + 3.75 = 6.75

Value 7 > 6.75

∴ 7 goals IS an outlier

Mean Formula: x̄ = Σfx ÷ Σf

16 of 41

Box & Whisker Plots — Examples 1 & 2

AQA Chapter 1.4 · Comparing distributions using five-number summaries

Key Rule: Always compare BOTH a measure of average (median) AND a measure of spread (IQR) when comparing two distributions from box plots.

Statistic

Class A

Class B

Min

8outlier

25

LQ

56

37

Median

59

59

UQ

63

70

Max

65

80

Example 1 — Class Marks (n = 15)

FIVE-NUMBER SUMMARY

Class A — IQR

7

Class B — IQR

33

Same median (59) — but Class B's IQR is 33 vs 7 : Class B marks are far more spread out. Class A has an outlier at 8 (below LQ − 1.5 × IQR = 45.5).

Example 2 — Cricket Scores (n = 11)

PLAYER COMPARISON

Ben

Min 0

LQ 6

Median 29

UQ 47

Max 91

Ben — IQR

41

Ben — Mean / SD

33.0  /  29.5

Sanjay

Min 14

LQ 25

Median 29

UQ 49

Max 61

Sanjay — IQR

24

Sanjay — Mean / SD

35.6  /  14.8

Pick

Sanjay

— higher mean (35.6 > 33.0) AND far more consistent (SD 14.8 vs 29.5)

17 of 41

SKILLS

Exercise 1C — How to Draw & Compare Box Plots

HOW TO DRAW A BOX PLOT

1

Sort data in ascending order

2

Find the five-number summary : Min, LQ, Median, UQ, Max

3

For n 20 — exclude median when finding LQ and UQ

4

Draw a number line with a suitable scale

5

Draw a box from LQ to UQ

6

Mark the median inside the box with a vertical line

7

Draw whiskers from box to min and max

8

Mark outliers as separate × points

OUTLIER FENCES

Lower fence =

LQ − 1.5 × IQR

Upper fence =

UQ + 1.5 × IQR

Any value outside these fences is an outlier

COMPARING TWO DISTRIBUTIONS

Always comment on median — measure of average

Always comment on IQR — measure of spread

WORKED EXAMPLE — BEN VS SANJAY (CRICKET SCORES, N = 11)

COMPARISON CONCLUSION

Median: Both players have the same median of 29 — similar average performance.

IQR: Sanjay IQR = 24 < Ben IQR = 41 — Sanjay is more consistent (less spread).

Always compare BOTH median (average) AND IQR (spread) in your answer

Min

LQ

Median

UQ

Max

0

6

29

47

91

Ben

SORTED DATA

0, 5, 6, 12, 20, 29 , 34, 43, 47, 75, 91

LQ = 3rd value = 6  |  UQ = 9th value = 47

IQR = 47 − 6 = 41

Outlier check: UQ + 1.5 × 41 = 47 + 61.5 = 108.5

Max = 91 < 108.5 No outliers

Min

LQ

Median

UQ

Max

14

25

29

49

61

Sanjay

SORTED DATA

14, 18, 25, 27, 28, 29 , 37, 46, 49, 58, 61

LQ = 3rd value = 25  |  UQ = 9th value = 49

IQR = 49 − 25 = 24

Outlier check: UQ + 1.5 × 24 = 49 + 36 = 85

Max = 61 < 85 No outliers

18 of 41

Exercise 1C — Questions

ATTEMPT ALL QUESTIONS BEFORE CHECKING ANSWERS

Age

17

18

19

20

21

Female

6

9

11

7

0

Male

13

11

8

5

9

Q1 — APPRENTICE AGES

Compare distributions for female and male apprentices (ages 17–21)

Compare the modal age , range , median age and mean age for each group. What do the differences tell you?

Q2 — CRICKET SCORES

Ben vs Sanjay — draw box plots and compare

BEN'S SCORES (N=11)

75

43

12

6

0

20

34

47

29

5

91

SANJAY'S SCORES (N=11)

28

49

27

61

29

14

58

37

25

18

46

Draw two box plots on the same scale. Compare mean and standard deviation . Who would you pick for the team and why?

Q3 — SUPERMARKET WAGES

Find median, IQR and check for outliers

£21,000

× 20 workers

£24,000

× 70 workers

£70,000

× 1 worker

Find the median and IQR . Show £70,000 is an outlier using UQ + 1.5 × IQR . Comment on the manager's claim that the average wage is £23,500 .

Q4 — BURGLARY RATES

Compare Badley and Lootham — find mean, SD, median, IQR and draw box plots

BADLEY (BURGLARIES/MONTH)

8

12

15

17

18

20

22

25

28

32

35

40

LOOTHAM (BURGLARIES/MONTH)

14

16

18

19

20

21

22

23

24

25

26

28

Find mean, SD, median and IQR for each town. Draw box plots and compare. Use statistical evidence to argue which town is safer and why.

19 of 41

Exercise 1C — Answers

MODEL ANSWERS

Q1

Apprentice Ages — Comparing Distributions

FEMALE (17–21)

Mode: 19

Range: 4

Median: 19

Mean: 18.9

MALE (17–21)

Mode: 17

Range: 4

Median: 18

Mean: 18.5

CONCLUSION

Males are slightly younger on average — all three averages are lower for males. Both groups have equal spread (range = 4 for both).

Q2

Cricket Scores — Ben vs Sanjay

Player

Mean

Std Dev (σ)

Verdict

Ben

33.0

29.5

High spread — inconsistent

Sanjay

35.6

14.8

Higher mean, far more consistent

PICK SANJAY

Higher mean score (35.6 vs 33.0) AND much lower standard deviation (14.8 vs 29.5) — better average and far more consistent performance.

Q3

Supermarket Wages — Outlier Check

FIVE-NUMBER SUMMARY

Median: £24,000

LQ: £21,000

UQ: £24,000

IQR: £3,000

OUTLIER FENCE CALCULATION

UQ + 1.5 × IQR = £24,000 + 1.5 × £3,000 = £28,500

Since £70,000 > £28,500 £70,000 IS an outlier

MANAGER'S CLAIM

Manager's £23,500 is the mean — distorted by the £70,000 outlier. Median £24,000 is more representative of typical wages.

Q4

Burglary Rates — Badley vs Lootham

Badley

Mean: 20.8

SD: 11.5

Median: 17

IQR: 22

Lootham

Mean: 20.3

SD: 9.0

Median: 20

IQR: 14

KYLIE — FEWER ACCIDENTS

Badley median = 17 < Lootham median = 20 Lower typical burglary rate.

WINSTON — MORE CONSISTENT

Lootham SD = 9.0 < Badley SD = 11.5 More predictable crime levels.

BOTH ARGUMENTS VALID

Depends on which measure you prioritise — median (average) or standard deviation (consistency). Always justify your choice with statistical evidence.

20 of 41

1D: Cumulative Frequency Graphs

Building and reading CF graphs to find median, quartiles and percentiles from grouped data

Key Rules

Plot at

upper class boundary

Median at

n ÷ 2

LQ at

n ÷ 4

UQ at

3n ÷ 4

Percentile:

p% × n

Always title & label axes!

Build

CF Table

Hours worked (x)

Frequency (f)

Upper boundary

0 < x ≤ 4

3

4

4 < x ≤ 8

29

8

8 < x ≤ 12

43

12

12 < x ≤ 16

10

16

16 < x ≤ 20

6

20

20 < x ≤ 24

1

24

Total

92

Step 1 — Frequency Table

Year 12 students' part-time hours per week  ( n = 92 )

1

Add frequencies cumulatively row by row

2

Always plot CF against the

upper class boundary

3

Draw a smooth S-curve through the plotted points

Upper boundary (x ≤)

Cumulative Frequency

Calculation

4

3

3

8

32

3 + 29

12

75

32 + 43

16

85

75 + 10

20

91

85 + 6

24

92

91 + 1

Step 2 — Cumulative Frequency Table

Plot each pair (upper boundary, CF) on your graph axes

Reading off the graph (n = 92)

Median: 46th value

 9.3 hrs

LQ: 23rd value

 7 hrs

UQ: 69th value

 11.5 hrs

IQR: 11.5 − 7

 4.5 hrs

90th percentile: 90% × 92 = 82.8th

 15 hrs

21 of 41

Exercise 1D — Questions & Answers

AQA CHAPTER 1D

Mass m (kg)

Frequency

Cum. Freq.

Upper Boundary

0 < m ≤ 100

11

11

100

100 < m ≤ 150

31

42

150

150 < m ≤ 200

16

58

200

200 < m ≤ 250

28

86

250

250 < m ≤ 300

19

105

300

300 < m ≤ 400

5

110

400

Total

110

Q1

Strawberry Farmer — Daily Mass Picked (n = 110 days)

A farmer records the mass of strawberries picked each day over 110 days. Use the frequency table to draw a cumulative frequency graph and find key statistics.

TASKS

(a)

State the modal class.

(b)

Draw a cumulative frequency graph (plot at upper class boundary).

(c)

Find the median, LQ, UQ, IQR, 40th and 80th percentiles.

(d)

Calculate the mean and standard deviation using midpoints.

ANSWERS

(A) MODAL CLASS

200 < m 250

(C) MEDIAN (55TH VALUE)

165 kg

LQ (28TH VALUE)

130 kg

UQ (83RD VALUE)

230 kg

IQR

100 kg

40TH PERCENTILE (44TH)

155 kg

80TH PERCENTILE (88TH)

250 kg

(D) MEAN & SD

185 kg, σ ≈ 72 kg

Attendance

Frequency

Cum. Freq.

Upper Boundary

Under 10,000

1

1

10,000

10,000 – 20,000

2

3

20,000

20,000 – 30,000

10

13

30,000

30,000 – 40,000

19

32

40,000

40,000 – 60,000

5

37

60,000

60,000 – 80,000

1

38

80,000

Total

38

Q2

Football Club Attendance — Season (n = 38 matches)

A football club records attendance at each of its 38 home matches. Use the grouped frequency table to draw a cumulative frequency graph and compare with last year's data (mean = 38,000).

TASKS

(a)

State the modal class.

(b)

Draw a cumulative frequency graph.

(c)

Find the median and IQR from your graph.

(d)

Estimate how many matches had attendance over 50,000. Is this more than 5% of matches?

(e)

Calculate the mean and standard deviation.

(f)

Compare this year's data with last year (mean = 38,000).

ANSWERS

(A) MODAL CLASS

30,000 – 40,000

(C) MEDIAN (19TH VALUE)

33,000

IQR

10,000

(D) OVER 50,000

2 matches

5% OF 38 = 1.9

Yes — just above 5%

(E) MEAN & SD

33,500, σ ≈ 10,000

(f) Comparison: Mean this year (33,500) is lower than last year (38,000), but SD is similar — attendances are slightly lower on average but equally spread. The lower mean may reflect fewer high-attendance fixtures.

22 of 41

1E: Histograms — Key Concepts & Frequency Density

AQA Mathematical Studies · Chapter 1E · School Traffic Survey Example

Tom's Mistake: Tom used HEIGHT to represent frequency —this is misleading when class widths are unequal. Always use AREA (FD × width) to represent frequency.

When to use histograms: Use histograms for continuous data with unequal class widths. No gaps between bars. Plot FD on y-axis, class boundaries on x-axis.

Activity 16: Verify each bar's area equals its frequency. Estimate the number of vehicles travelling at less than 15 mph using the histogram.

Frequency Density = Frequency ÷ Class Width

Key Formula

Area = Frequency

NOT Height

Frequency = FD × Width

Reading a Histogram

Speed (mph)

Frequency

Class Width

Freq. Density

Area Check

0 < x ≤ 10

1

10

0.1

0.1 × 10 = 1

10 < x ≤ 20

8

10

0.8

0.8 × 10 = 8

20 < x ≤ 26

6

6

1.0

1.0 × 6 = 6

26 < x ≤ 30

28

4

7.0

7.0 × 4 = 28

30 < x ≤ 32

14

2

7.0

7.0 × 2 = 14

32 < x ≤ 34

3

2

1.5

1.5 × 2 = 3

34 < x ≤ 36

0

2

0

0 × 2 = 0

36 < x ≤ 38

2

2

1.0

1.0 × 2 = 2

38 < x ≤ 40

1

2

0.5

0.5 × 2 = 1

Histogram: Frequency Density vs Speed — Total = 65 vehicles

23 of 41

Exercise 1E — Q1 & Q2

AQA CHAPTER 1E · HISTOGRAMS & FREQUENCY DENSITY

Key Formula:

Frequency Density (FD) = Frequency ÷ Class Width

Given

Complete this

FD value

Q1 — Supermarket Spending (£P)

TASKS

(a) Copy and complete the table — state any assumptions you make.

(b) Draw a histogram to represent this data.

* Assumption needed: P < 25 starts at 0; P 250 ends at 350 (width = 100)

Class (£P)

Freq.

LCB

UCB

Width

FD = f ÷ w

P < 25

75

?

25

25

3.00

25 ≤ P < 50

97

25

50

?

3.88

50 ≤ P < 100

165

50

100

50

?

100 ≤ P < 150

86

100

150

?

1.72

150 ≤ P < 250

23

150

250

100

?

P ≥ 250

8

250

350*

100*

?

Q2 — UK Population by Age (June 2013)

TASKS

(a) Copy and complete the table — state any assumptions you make.

(b) Draw a histogram to represent the UK age distribution.

* Assumption needed: 90+ ends at 100 (width = 10) — state this clearly

Age Group

Freq. (000s)

LCB

UCB

Width

FD = f ÷ w

0 – 4

4,014

0

5

?

802.8

5 – 15

11,179

5

16

11

?

16 – 44

21,453

16

45

?

739.8

45 – 64

16,328

45

65

20

?

65 – 74

6,031

65

75

?

603.1

75 – 89

4,574

75

90

15

?

90+

527

90

100*

10*

?

· For open-ended classes, you must state your assumption about the boundary

n = 454 customers

Frequency in thousands · Source: ONS 2014

24 of 41

EXERCISE 1E · Q1 · HISTOGRAMS & FREQUENCY DENSITY

Q1 — Supermarket Spending: Completed Table & Histogram

KEY FORMULA

FD = Frequency ÷ Class Width

Class (£P)

Freq

LCB

UCB

Width

FD

P < 25

75

0

25

25

3.00

25 ≤ P < 50

97

25

50

25

3.88

50 ≤ P < 100

165

50

100

50

3.30

100 ≤ P < 150

86

100

150

50

1.72

150 ≤ P < 250

23

150

250

100

0.23

P ≥ 250 *

8

250

350*

100*

0.08

Total

454

* Assumption: P 250 class ends at £350 (width = 100). This must be stated clearly in your answer.

KEY POINTS

Gold values = answers to fill in. Teal = given in question.

FD = Frequency ÷ Width: e.g. 165 ÷ 50 = 3.30

Area of each bar = Frequency (not height). Bars have no gaps .

Highest FD bar is 50P<100 (FD=3.30) — most customers spend £50–£100.

HISTOGRAM — SUPERMARKET SPENDING (N = 454 CUSTOMERS)

25 of 41

EXERCISE 1E · Q2 · HISTOGRAMS & FREQUENCY DENSITY

Q2 — UK Population Age Distribution (June 2013)

KEY FORMULA

FD = Frequency (000s) ÷ Class Width

Age Group

Freq (000s)

Width

FD

0 – 4

4,014

5

802.8

5 – 15

11,179

11

1016.3

16 – 44

21,453

29

739.8

45 – 64

16,328

20

816.4

65 – 74

6,031

10

603.1

75 – 89

4,574

15

304.9

90+ *

527

10*

52.7

Total

64,106

* Assumption: 90+ group ends at age 100 (width = 10). State this clearly in your answer.

KEY POINTS

5–15 group : width = 16 − 5 = 11 FD = 11,179 ÷ 11 = 1016.3

16–44 group : width = 45 − 16 = 29 FD = 21,453 ÷ 29 = 739.8

Tallest bar: 5–15 (FD=1016.3) — large group due to 11-year span.

Unequal widths must use FD , not frequency, for bar height.

HISTOGRAM — UK POPULATION BY AGE GROUP (SOURCE: ONS 2014)

26 of 41

Exercise 1E — Q3, Q4 & Q6

Frequency Density Tables & Histograms  |  AQA Chapter 1E

IQ Interval

Freq

Width

FD

85 < x ≤ 95

21

10

?

95 < x ≤ 100

76

5

?

100 < x ≤ 105

137

5

?

105 < x ≤ 110

129

5

?

110 < x ≤ 115

108

5

?

115 < x ≤ 125

83

10

?

125 < x ≤ 135

16

10

?

Total

570

Q3 — IQ SCORES (N = 570)

Gainsby College student IQ distribution

TASKS

(a)(i) Calculate estimate for mean IQ and standard deviation.

(a)(ii) Last year: mean = 106, SD = 7.4. Write two comparison statements.

(b)(i) Draw a histogram using frequency density.

(b)(ii) Use histogram to estimate % with IQ > 106.

Delay (min)

Flights

Width

FD

Early (−15 to 0)

12,618

15

?

1 – 15 late

15

?

16 – 30 late

2,460

15

?

31 – 60 late

1,589

30

?

61 – 180 late

1,029

120

?

181 – 360 late

216

180

?

> 360 late

30

?

Q4 — GATWICK FLIGHT DELAYS

Civil Aviation Authority data — Will's histogram is incorrect

WILL'S ERROR

Will used bar height = frequency instead of area = frequency . Class widths are unequal, so frequency density must be used.

TASKS

(a) Give reasons why Will's histogram is incorrect.

(b) Draw a correct histogram using frequency density.

Width (mm)

Doors

Width

FD

Less than 750

0

0

750 – 790

24

45*

?

800 – 890

146

90

?

900 – 990

124

90

?

1000 – 1090

68

90

?

1100 and over

21

?

Total

383

Q6 — DOOR WIDTHS (N = 383)

Local authority doors, rounded to nearest 10 mm

BOUNDARY ASSUMPTION *

750–790 group: widths rounded to nearest 10 mm, so boundaries are 745 – 795 →width = 50 mm (not 40). State this assumption clearly.

TASKS

(a) Estimate number of doors not meeting official guideline (≥ 775 mm).

(b) Estimate % of doors the LA wants to widen (target: ≥ 850 mm).

Guidelines: Official ≥ 775 mm  |  LA target ≥ 850 mm

FD = Frequency ÷ Class Width

Area under histogram = Frequency

Unequal class widths must use FD (not frequency) for histogram height

Always state assumptions for open-ended or ambiguous class boundaries

27 of 41

EXERCISE 1E · Q3 · IQ SCORES (N = 570)

Q3 — Gainsby College IQ Scores: FD Table, Histogram & Statistics

IQ Interval

Freq

Mid

Width

FD

85 < x ≤ 95

21

90

10

2.1

95 < x ≤ 100

76

97.5

5

15.2

100 < x ≤ 105

137

102.5

5

27.4

105 < x ≤ 110

129

107.5

5

25.8

110 < x ≤ 115

108

112.5

5

21.6

115 < x ≤ 125

83

120

10

8.3

125 < x ≤ 135

16

130

10

1.6

Total

570

MEAN IQ

103.4

Σfx ÷ 570 = 58,965 ÷ 570

STD DEV

8.0

Use σₙ₋₁ on calculator

COMPARISON WITH LAST YEAR (MEAN=106, SD=7.4)

This year's mean ( 103.4 ) is lower than last year's (106)— students scored lower on average.

This year's SD (8.0 ) is slightly higher than last year's (7.4) — scores are more spread out.

% with IQ > 106: Bars from 105–135 = 129+108+83+16 = 336. But 106 is 1/5 into the 105–110 bar, so approx 4/5 ×129 + 108 + 83 + 16 310 ÷ 570 54%

HISTOGRAM — IQ SCORE DISTRIBUTION (N = 570 STUDENTS) · VERTICAL LINE AT IQ = 106

28 of 41

EXERCISE 1E · Q4 · GATWICK FLIGHT DELAYS — WILL'S ERROR & CORRECT HISTOGRAM

Q4 — Gatwick Delays: Why Will Was Wrong & Correct FD Histogram

Delay (min)

Flights

Width

FD = f÷w

Early (−15 to 0)

12,618

15

841.2

1–15 late

not given

15

16–30 late

2,460

15

164.0

31–60 late

1,589

30

53.0

61–180 late

1,029

120

8.6

181–360 late

216

180

1.2

>360 late

30

open*

excluded

(A) WHY WILL'S HISTOGRAM WAS WRONG

Error 1: Will used frequency as bar height — but class widths are unequal, so this misrepresents the data.

Error 2: In a histogram, area = frequency , not height. Height must be frequency density (FD = f ÷ width).

Example: The 61–180 min bar (width=120) would appear far too tall if frequency (1,029) is used as height instead of FD (8.6).

* Note: The >360 min class has no upper boundary given — it cannot be plotted on a histogram without making an assumption. State this clearly and exclude it, or assume a boundary (e.g., 540 min).

CORRECT HISTOGRAM — GATWICK FLIGHT DELAYS (FD ON Y-AXIS, AREA = FREQUENCY)

29 of 41

EXERCISE 1E · Q6 · DOOR WIDTHS (N = 383)

Q6 — Door Widths: FD Table, Histogram & Guideline Estimates

Width (mm)

Doors

Boundaries

Width

FD

Less than 750

0

0

750–790 *

24

745–795

50

0.48

800–890

146

795–895

90

1.62

900–990

124

895–995

90

1.38

1000–1090

68

995–1095

90

0.76

1100 and over *

21

1095–1185*

90*

0.23

Total

383

* Key Assumption: 750–790 group: widths rounded to nearest 10 mm boundaries are 745–795 (width = 50 mm, not 40). State this clearly. 1100+ assumed to end at 1185 (width = 90).

(A) DOORS < 775 MM (OFFICIAL GUIDELINE)

14 doors

745–795 band: 30/50 of 24 doors below 775

= 0.6 × 24 = 14.4 14 doors

(B) % OF DOORS < 850 MM (LA TARGET)

29.5%

All 745–795 (24) + 55/90 of 795–895 (146)

= 24 + 89.2 = 113.2 113÷383 = 29.5%

HISTOGRAM — DOOR WIDTHS (N = 383) · DASHED LINES AT 775 MM AND 850 MM GUIDELINES

30 of 41

1F: Choosing Statistical Methods

SUMMARY

Key Principle: Always think carefully about whether your chosen methods are appropriate and help communicate your message clearly to the audience.

AVERAGES & SPREAD

Mean & Standard Deviation

Uses all data values; sensitive to extreme values and outliers.

Median & IQR

Not affected by extreme values; does not use all data in calculation.

Mode & Range

Easy to find; range is heavily affected by extreme values.

CHARTS & DIAGRAMS

Bar Charts, Pictograms & Pie Charts

Qualitative & ungrouped discrete data; pie shows proportions, not frequencies.

Stem-and-Leaf Diagrams

Quantitative data; organises raw values, enables mode, median and quartiles.

Box & Whisker Plots

Shows min, max, median and quartiles; does not display individual frequencies.

Cumulative Frequency Diagrams

Continuous & grouped discrete; plot at upper class boundary to find quartiles.

Histograms

Continuous & grouped discrete; area = frequency; use FD for unequal class widths.

31 of 41

Consolidation Exercise 1 — Q1, Q2 & Q3

CHAPTER 1 · SAMPLING & DATA

Q1

SAMPLING METHODS

(a)(i) Random sample:

Every member of the population has an equal chance of being selected.

(a)(ii) Representative:

A sample that reflects the characteristics of the whole population.

(b)(i) Cluster:

e.g. Select 3 schools at random, then survey all students in those schools.

(b)(ii) Quota:

e.g. Interview 10 males and 10 females aged 18–25 in a shopping centre.

Group

Count

Calculation

Sample

Year 1 Male

17

17/90 × 20 = 3.8

4

Year 1 Female

12

12/90 × 20 = 2.7

3

Year 2 Male

19

19/90 × 20 = 4.2

4

Year 2 Female

11

11/90 × 20 = 2.4

2

Year 3 Male

23

23/90 × 20 = 5.1

5

Year 3 Female

8

8/90 × 20 = 1.8

2

Total

90

20 ✓

Q2

STRATIFIED SAMPLE — ENGINEERING COURSE (N = 90, SAMPLE = 20)

(b) Suggest two ways to select students from each group (e.g. random number table, lottery method).

Q3

PLANT HEIGHTS (CM) — OUTDOORS VS INDOORS

OUTDOORS (N = 22)

27, 32, 16, 10, 18, 13, 25, 23, 9, 11,

17, 21, 20, 17, 18, 29, 29, 15, 22, 11, 26, 31

INDOORS (N = 20)

17, 22, 23, 27, 35, 18, 27, 14, 32, 10,

27, 28, 30, 34, 12, 15, 22, 11, 26, 31

(a)(i) Draw a back-to-back stem-and-leaf diagram for both groups.

(a)(ii) Write down the mode for each group.

(a)(iii) Work out the range for each group.

(b)(i) Draw a box and whisker plot for each group.

(b)(ii) Describe and compare the two distributions.

32 of 41

Consolidation Exercise 1 — Q4 & Q5

Data Tables & Tasks

Key Reminder — Q5: Class widths are unequal. You must use Frequency Density = Frequency ÷ Class Width on the vertical axis of your histogram. Do not plot raw frequencies.

Q4 — French Test Marks

A Level French, n = 64 students

Mark Range

Frequency

Cumulative Freq.

1 – 10

0

0

11 – 20

8

8

21 – 30

11

19

31 – 40

23

42

41 – 50

17

59

51 – 60

5

64

Total

64

TASKS

a(i)

Draw a cumulative frequency graph for the data.

a(ii)

Find the 40th percentile. What does this value tell you?

a(iii)

The pass mark is 40%. Estimate the number of students who passed.

b

Draw a box and whisker plot to illustrate the distribution of marks.

Q5 — Clothes Spending

Julie's survey, n = 88 students

Amount Spent (£P)

Frequency

Class Width

P < 20

8

20

20 ≤ P < 30

19

10

30 ≤ P < 40

27

10

40 ≤ P < 60

18

20

60 ≤ P < 100

14

40

P ≥ 100

2

Total

88

TASKS

a

Julie says "My data is continuous, secondary data." Is she correct? Explain your answer fully.

b

Draw a histogram for Julie's data. Remember: unequal class widths require frequency density.

c

Find an estimate for the mean amount spent on clothes and the standard deviation.

33 of 41

Consolidation Exercise 1 — Q6, Q7 & Q8

AQA MATHEMATICAL STUDIES

Q6 — PETROL & DIESEL CONSUMPTION GRAPHS

CONTEXT

UK consumption of petrol and diesel, 2000–2013. Two graphs are provided:

TANYA'S GRAPH

No scale on y-axis; uses misleading pictogram-style bars — bar widths vary, making comparison unreliable.

OLIVER'S GRAPH

Proper numerical scale on y-axis; consistent bar widths; clearly labelled axes.

TASKS

(a) List the ways in which each diagram can be improved.

(b) Describe how the consumption of petrol and diesel has changed over this time period.

Consider: missing title, missing units, no key, misleading scale, pictogram distortion.

Q7 — BLACKPOOL SCHOOLS SAMPLING

BLACKPOOL LOCAL AUTHORITY

Total: 63 schools and colleges

39

Primary

Schools

17

Secondary

Schools

7

Colleges

(16–18)

PROPOSED METHOD

Choose one school/college from each group at random, then interview two teachers from each selected institution.

TASKS

(a) Give two reasons why this is not a good sample.

(b) Describe a better sampling method for this situation.

Hint: consider sample size, proportional representation, and whether two teachers per school is sufficient.

Q8 — BOX OFFICE TAKINGS (BFI 2014) — HISTOGRAM

Takings T (£m)

Frequency

Class Width

0 < T < 0.1

433

0.1

0.1 ≤ T < 1

133

0.9

1 ≤ T < 5

70

4

5 ≤ T < 10

28

5

10 ≤ T < 20

20

10

20 ≤ T < 30

6

10

30 ≤ T < 50

8

20

Total

698

DATA TABLE — 698 FILMS

Unequal class widths — use

frequency density

for histogram (fd = freq

÷ class width)

TASKS

(a) Write down the modal class .

(b) Estimate the mean and standard deviation .

(c) Estimate the median and IQR . Describe any problems.

(d) Write three sentences explaining what the data shows.

Source: BFI Statistical Yearbook 2014

34 of 41

Consolidation Exercise 1 — Q9 & Q10

AQA Mathematical Studies

Q9 — HOSPITAL WAITING TIMES

Waiting Time

% of Patients

Less than 5 weeks

2%

5–9 weeks

17%

10–15 weeks

26%

16–17 weeks

38%

18 weeks (target boundary)

12%

19 weeks

4%

20 weeks

1%

More than 20 weeks

0%

Non-emergency treatment waiting times (Target: no more than 18 weeks)

Highlighted row = 18-week target boundary

TASKS

Comment on the hospital's performance against the 18-week target.

Use statistical measures and/or diagrams to support your comments.

Hint: Consider cumulative frequencies, median, and what % waited 18 weeks.

Q10 — UK GENDER GAP: YEARS OF LIFE LOST

Age Group

F Yrs Lost

M Yrs Lost

F Pop

M Pop

F Deaths

M Deaths

20–24

5,544.0

4,582.0

1,774,400

1,829,400

90

79

30–34

12,873.3

13,620.6

1,850,300

1,831,800

249

282

50–54

55,094.0

65,072.7

1,824,700

1,792,800

1,690

2,191

70–74

89,900.0

111,577.5

1,106,900

998,900

5,800

8,265

Premature death data by age group — UK Gender Gap Report (Oct 2014). UK ranked 26th (WEF).

TASKS

(a) How does the mean number of years lost by females in each age group compare with males?

(b) Write down two more questions you could ask based on this data.

Hint: Calculate mean years lost per death for each gender and age group.

Source: Health & Social Care Information Centre (HSCIC)

35 of 41

Q11 — Gatwick vs Heathrow: CAA Passenger Survey 2013

Consolidation Exercise 1

YOUR TASK

Use the data in all three tables to compare the use of Gatwick and Heathrow by passengers who travel by air. Consider group sizes, trip lengths, and age profiles. Use statistical measures (mean, median, percentages) and diagrams to support your comparisons.

Extension: Use the internet to find and compare data for other UK airports (e.g. Manchester, Birmingham, Edinburgh).

SOURCE

Civil Aviation Authority (CAA) Passenger Survey Report 2013.

Gatwick: 32,402,000 passengers  |  Heathrow: 45,744,000 passengers

No. in Group

Gatwick %

Heathrow %

1

36.8

63.3

2

36.1

23.9

3

4.9

4.5

4

12.8

4.1

5

3.7

1.2

6+

5.8

3.1

Total passengers (000s)

32,402

45,744

TABLE 1 — TRAVELLING GROUP SIZE

Length (x days)

Gatwick %

Heathrow %

x < 1

2.7

4.7

1 ≤ x < 3

11.4

13.1

3 ≤ x < 6

30.4

25.1

6 ≤ x < 8

24.3

12.7

8 ≤ x < 15

23.4

21.8

x ≥ 15

7.8

22.7

Total

100%

100%

TABLE 2 — TRIP LENGTH (DAYS)

Age Group

Gatwick %

Heathrow %

10 or under

3.2

0.9

11–15

3.4

1.6

16–19

4.9

4.2

20–24

10.5

9.8

25–34

18.5

24.0

35–44

16.0

20.3

45–54

17.9

18.5

55–59

7.6

7.3

60–64

8.3

6.1

65–74

8.0

5.9

Over 74

1.7

1.4

Total

100%

100%

TABLE 3 — AGE DISTRIBUTION

36 of 41

Exercise 1F — Activity 17 & 18

RAINFALL DATA

England & Wales — Hadley Centre Data

Year

Rainfall (mm)

Note

1770

1079.4

Highest

1771

792.9

Lowest

1772

1031.8

1773

1033.8

1774

994.7

1775

1012.0

1776

847.4

1777

860.3

1778

887.9

1779

900.5

TABLE 1 — 1770 TO 1779 (MM)

Year

Rainfall (mm)

Note

2000

1232.5

Highest

2001

970.0

2002

1117.8

2003

761.4

Lowest

2004

973.6

2005

825.1

2006

904.8

2007

1022.7

2008

1089.6

2009

977.1

TABLE 2 — 2000 TO 2009 (MM)

ACTIVITY 18 — STUDENTS' METHODS

Ahmed

Bar chart comparing monthly rainfall for two years a century apart (1766 vs 2014)

Kayleigh

Time series graph of annual rainfall 1766–2010 (long-term trend)

Cilla

Mean & SD of monthly rainfall for selected years

Kirsty

Pie charts for first and last years (1766 & 2014)

Tim

Median & IQR for 30-year groups (1771–2010)

FOR EACH STUDENT, EVALUATE:

(a)

Was the method appropriate?

(b)

Describe any patterns or changes found

(c)

How could their work be improved?

(d)

What other methods could they have used?

ANNUAL RAINFALL COMPARISON — 1770S VS 2000S (MM)

37 of 41

Activity 18 — Students' Statistical Methods

EXERCISE 1F

FIVE STUDENTS INVESTIGATED HADLEY CENTRE RAINFALL DATA (1766 ONWARDS):

Ahmed

Bar chart — two years compared

Kayleigh

Time series — 1766–2010

Cilla

Mean & SD per year

Kirsty

Pie charts — 1766 & 2014

EVALUATE EACH STUDENT'S WORK:

(a) Was the method appropriate for investigating rainfall changes over time?

(b) Describe any patterns or changes found in the data.

(c) How could their work be improved?

(d) What other statistical methods could they have used?

CILLA — MEAN & SD OF MONTHLY RAINFALL (MM)

Year

Mean (mm)

SD (mm)

1766

63.0

33.8

1767

78.4

32.0

1768

103.9

38.7

1769

79.2

27.9

1770

90.0

41.0

2010

68.5

25.1

2011

65.6

28.5

2012

103.7

46.9

2013

76.4

35.4

2014

92.1

45.2

TIM — MEDIAN & IQR OF ANNUAL RAINFALL BY 30-YEAR GROUP

Period

Median (mm)

IQR (mm)

1771–1800

888.35

199.6

1801–1830

903.55

157.35

1831–1860

873.95

169.3

1861–1890

912.45

183.95

1891–1920

907.45

155.45

1921–1950

925.4

165.65

1951–1980

904.5

206.65

1981–2010

971.8

170.25

38 of 41

Chapter Summary — Key Points & Review

CHAPTER 1

Data Types

Know the meaning of key terms for classifying data correctly before analysis.

Qualitative

Quantitative

Discrete

Continuous

Primary

Secondary

Sampling Methods

Appreciate strengths and limitations of each method; larger samples improve accuracy but cost more time and money.

Random

Cluster

Stratified

Quota

Stratified: (group size ÷ total) × sample size

Numerical Measures

Represent data numerically from raw data, frequency tables, or statistical diagrams.

Mean

Median

Mode

Quartiles

IQR

Std Dev

Mean from grouped data: use midpoints of each class. Use σₙ₋₁ on calculator.

Statistical Diagrams

Represent data appropriately and interpret diagrams to reach valid conclusions.

Histograms

CF Graphs

Stem-and-Leaf

Box Plots

Key Reminders

Critical rules to avoid common errors in exams and coursework.

CF graphs: always plot at upper class boundary

Histograms: FD = frequency ÷ class width ; area = frequency

Percentiles: nth percentile at n% × total

Essential Formulae

Core formulae required for all Chapter 1 calculations and exam questions.

Stratified sample: (group ÷ total) × n

nth percentile: n% × Σf on CF graph

Grouped mean: Σ(midpoint × f) ÷ Σf

39 of 41

AQA MATHEMATICAL STUDIES — CHAPTER 1 COMPLETE

Chapter 1 Complete

Analysis of Data — from data types and sampling through to histograms and consolidation

TOPICS MASTERED

Cumulative Frequency

Graphs & percentiles

Histograms

Frequency density

Statistical Methods

Choosing appropriately

Consolidation Exercise

All Chapter 1 topics

NEXT STEPS

Practise past AQA exam questions on data analysis

Use your calculator for mean and SD from grouped data

Review the key formulae card before your next lesson

WELL

DONE!

40 of 41

1E: Histograms — Frequency Density

AQA MATHEMATICAL STUDIES

Speed (mph)

Freq (f)

Width (w)

FD = f ÷ w

0 < x ≤ 10

1

10

0.10

10 < x ≤ 20

8

10

0.80

20 < x ≤ 26

6

6

1.00

26 < x ≤ 30

28

4

7.00

30 < x ≤ 32

14

2

7.00

32 < x ≤ 34

3

2

1.50

34 < x ≤ 36

0

2

0.00

36 < x ≤ 38

2

2

1.00

38 < x ≤ 40

1

2

0.50

SCHOOL TRAFFIC SURVEY — SPEED DATA (N = 65 VEHICLES)

Highlighted rows = peak FD (speed limit zone 26–32 mph)

HISTOGRAM — FREQUENCY DENSITY VS SPEED

KEY FORMULA

FD = Frequency ÷ Class Width

To read back: Frequency = FD × Class Width

GOLDEN RULE

AREA represents frequency — not height

Use for continuous data with unequal class widths. No gaps between bars.

COMMON MISTAKE

Tom's diagram was misleading

He used HEIGHT not AREA to represent frequency — this is incorrect for unequal widths.

41 of 41

CHAPTER 1

Summary — Key Points & Review

Key Reminders:

Always compare average AND spread

Always check for outliers using fences

Plot CF at upper class boundary

Use FD for unequal class width histograms

Data Types

Qualitative — descriptive categories.

Discrete — exact countable values.

Continuous — any value in a range.

Sampling

Random, cluster, quota methods.

Stratified sample size per group:

(group ÷ total) × sample size

Averages

Mode — most frequent value.

Median — middle value when ordered.

Mean = Σx/n  or  Σfx/Σf

Spread & Outliers

Use σₙ₋₁ on calculator for SD.

IQR = UQ − LQ

Outlier: < LQ−1.5×IQR  or  > UQ+1.5×IQR

Box Plots

Five-number summary: Min, LQ, Median, UQ, Max.

Mark outliers with ×. Always compare median (average) and IQR (spread) when comparing two distributions.

Cumulative Frequency

Plot CF against upper class boundary . Draw smooth S-curve.

Median = n/2  ·  LQ = n/4  ·  UQ = 3n/4

Histograms

Used for continuous data with unequal class widths .

Area represents frequency — not height.

FD = Frequency ÷ Class width