1 of 29

M.A. EDUCATION    EDUCATIONAL MEASUREMENT AND EVALUATION

UNIT III

Test Construction & Standardization

PrinciplesItem TypesSteps of ConstructionItem Analysis

A Comprehensive Study Presentation

2 of 29

OVERVIEW

What This Presentation Covers

Four core themes of Unit III, built up from foundational principles to the statistical tools used to refine a test.

01

Principles of Test Construction & Standardization

What makes a test valid, reliable, objective and fair

02

Types of Test Items

Essay, short-answer and objective-type items compared

03

Steps of Test Construction

From planning and blueprint through to standardization

04

Item Analysis

Difficulty, discrimination and distracter analysis

Unit III | Test Construction & Standardization

2

3 of 29

1

Principles of Test Construction

& Standardization

Meaning of a test, test construction, standardization — and nine hallmarks of a good test

Unit III | Test Construction & Standardization

3

4 of 29

1.1 – 1.3

Three Key Terms, Clearly Defined

A Test

A systematic procedure for measuring a sample of behaviour — knowledge, skill, attitude or trait — under standard conditions, scored in a uniform, objective manner.

Test Construction

The entire scientific, sequential process — from deciding what to measure, to writing, trying out, analysing and finalising items — that produces a valid, reliable, usable instrument.

Standardization

Making a test uniform in content, administration, scoring and interpretation, so scores from different individuals, times and places are genuinely comparable.

1. Principles of Test Construction & Standardization

4

5 of 29

1.3

A Standardized Test Always Has…

Fixed Content

Selected scientifically through try-out and item analysis

Uniform Instructions

Same time limit, materials and verbal instructions for all

Objective Scoring

A fixed scoring key or rubric, free of examiner judgement

Established Norms

Age, grade or percentile norms for interpreting a score

Known Psychometrics

Reliability & validity coefficients reported in a test manual

1. Principles of Test Construction & Standardization

5

6 of 29

1.4 (A)–(C)

Characteristics of a Good Test

Validity

The test measures what it claims to measure — nothing more, nothing less. Content validity comes from a proper Blueprint; criterion and construct validity are established statistically.

Reliability

The test yields consistent scores on repeated administration, verified via test–retest, split-half, Kuder–Richardson or Cronbach's alpha.

Objectivity

Items and scoring are free from the scorer's personal bias — two independent examiners must arrive at the same score.

1. Principles of Test Construction & Standardization

6

7 of 29

1.4 (D)–(F)

Characteristics of a Good Test

Availability of Norms

Adequate, representative norms — from a large sample — let an individual's raw score be meaningfully interpreted.

Practicability

Economical in cost, time and effort; easy to administer, score and interpret, with clear instructions and durable material.

Comprehensiveness

Items adequately sample the entire content domain and every level of objective reflected in the Blueprint.

1. Principles of Test Construction & Standardization

7

8 of 29

1.4 (G)–(I)

Characteristics of a Good Test

Discriminating Power

The test differentiates effectively between high and low achievers — every item, and the test overall, must discriminate well.

Adequate Difficulty

Items are neither too easy nor too difficult; a balanced spread approximating a normal distribution is desirable.

Fairness

Free from cultural, linguistic, gender or regional bias, so no group is unduly favoured or disadvantaged.

1. Principles of Test Construction & Standardization

8

9 of 29

2

Types of Test Items

Essay, short-answer and objective-type items — their forms, merits and limitations

Unit III | Test Construction & Standardization

9

10 of 29

2.1

Essay Type Items

Items requiring the examinee to organise and present an answer in their own words — a paragraph to several pages — with freedom to select, organise and express content.

TWO SUB-TYPES

1

Extended-Response Essay

No restriction on length, organisation or content; assesses higher-order abilities — analysis, synthesis, evaluation.

2

Restricted / Short-Response Essay

Scope narrowed by specifying the form, length or content expected (e.g., ‘Explain in about 200 words…’).

2. Types of Test Items

10

11 of 29

2.1

Essay Type — Merits & Limitations

MERITS

Measures higher mental processes — organisation, analysis, synthesis, evaluation, creativity

Easy and quick to construct

Encourages original thinking; reduces guessing

Suitable for measuring writing ability itself

LIMITATIONS

Scoring is time-consuming, subjective, prone to halo-effect & examiner fatigue

Limited content sampling reduces validity/reliability

Encourages rote-memorised, bluffing or padded answers

2. Types of Test Items

11

12 of 29

2.2

Short-Answer Type Items

Items demanding a brief, precise answer — a word, phrase, sentence, number or formula. They occupy an intermediate position between essay and objective items.

MERITS

Wider content coverage than essay items in the same time

Scoring is more objective than essay items

Reduces guessing and bluffing to a great extent

Easy to construct; good for facts, terms, definitions, simple computations

LIMITATIONS

Not suitable for measuring organisation and expression ability

Ambiguous wording may permit more than one ‘correct’ answer

Encourages memorisation of isolated facts over deeper understanding

2. Types of Test Items

12

13 of 29

2.3

Objective Type Items — Common Forms

Items with a fixed, predetermined correct answer, permitting only one way of scoring.

Multiple Choice (MCQ)

A stem/question with one correct answer and 2–4 plausible distracters.

True – False

A statement to be judged correct or incorrect (alternate-response).

Matching Items

Two parallel lists — premises and responses — to be matched.

Completion / Fill-in-the-Blank

A statement with blank(s) to be filled with the exact missing word(s).

2. Types of Test Items

13

14 of 29

2.3

Objective Type — Merits & Limitations

MERITS

Highly objective, reliable and quick to score — even machine-scorable

Wide, representative content sampling improves content validity

Eliminates the bluffing/writing-skill factor; measures knowledge directly

Amenable to statistical item analysis (difficulty, discrimination, distracters)

LIMITATIONS

Difficult, time-consuming and skill-demanding to construct — especially good distracters

Cannot measure organisation, expression or writing ability

Permits guessing (correctable) and favours recognition over recall

2. Types of Test Items

14

15 of 29

2.4

Comparative Summary

Basis

Essay Type

Short-Answer Type

Objective Type

Length of answer

Long, extended

Word / phrase / sentence

Selection / one word

Content coverage

Very limited

Moderate

Wide / comprehensive

Ease of construction

Easiest

Easy

Most difficult

Objectivity of scoring

Low (subjective)

Moderate

Very high

Scoring time

Longest

Moderate

Shortest

Guessing factor

Almost nil

Low

Present (correctable)

Higher-order thinking

Best measured

Limited

Possible with careful design

2. Types of Test Items

15

16 of 29

3

Steps of Test Construction

& Standardization

With special reference to Achievement Tests — from Planning to a standardized Test Manual

Unit III | Test Construction & Standardization

16

17 of 29

1

Planning

The foundation stage — key decisions taken before a single item is written:

Purpose

Diagnostic, formative, summative, placement, selection or research

Content & Objectives

Syllabus and instructional objectives, stated behaviourally per Bloom's Taxonomy

Target Group

The age, grade or class for whom the test is meant

Item Types & Length

Type(s) of items to use, total number of items, and total time

Difficulty Level

The intended level of difficulty and overall length of the test

3. Steps of Test Construction & Standardization

17

18 of 29

2

Preparing the Design & Blueprint

The Blueprint is a three-dimensional table fixing the exact number of items per topic, per objective and per item-form.

Content / Unit

Knowledge

Understanding

Application

Skill

Total

Unit I

2 (1)

2 (1)

1 (1)

1 (1)

6

Unit II

2 (1)

3 (1)

2 (1)

1 (1)

8

Unit III

1 (1)

2 (1)

2 (1)

1 (1)

6

Total

5

7

5

3

20

WHY IT MATTERS

The Blueprint ensures content validity — every important topic and every level of objective gets fair, proportionate representation, preventing overemphasis on any single unit.

Figures = marks; brackets = number of items.

3. Steps of Test Construction & Standardization

18

19 of 29

3

Item Writing — Golden Rules

1

State each item clearly, precisely and in simple, unambiguous language

2

One item should test one idea / objective only

3

Avoid textbook language and clues that give away the answer (e.g., ‘always’, ‘never’)

4

Arrange items from easy to difficult; group items of the same form together

5

For MCQs: write a clear stem, keep options parallel and plausible, avoid overlap

6

Prepare more items than required (1.5–2×) — some will be rejected after try-out

7

Prepare the scoring key / marking scheme alongside the items

3. Steps of Test Construction & Standardization

19

20 of 29

4–5

Try-out & Item Analysis

Step 4 — Preliminary Try-out

Administer the draft test to a small, representative sample

Follow intended instructions & time limit exactly

Collect and score try-out data for analysis

Step 5 — Item Analysis

Statistical scrutiny of try-out responses

Compute Difficulty Index, Discrimination Index & distracter effectiveness

Decide to retain, revise or reject each item (see Section 4)

3. Steps of Test Construction & Standardization

20

21 of 29

6–7

Final Draft & Standardization

STEP 6 — FINAL DRAFT

Select the best items per the original Blueprint proportions

Finalise instructions, time limit and scoring key

Print the test booklet and response sheet

STEP 7 — STANDARDIZATION (LARGE-SAMPLE ADMINISTRATION)

Reliability

Test–retest, split-half, KR / Cronbach's alpha

Validity

Content, criterion-related, and construct validity

Norms

Percentile, grade, age, or standard-score norms

Test Manual

Administration, scoring, norms & psychometric data

3. Steps of Test Construction & Standardization

21

22 of 29

4

Improving Quality Through

Item Analysis

Difficulty, discrimination and distracter analysis — the statistical tools that refine a test

Unit III | Test Construction & Standardization

22

23 of 29

4.1 – 4.2

Meaning, Purpose & Procedure

Item analysis is the statistical study of examinees' responses to each try-out item, used to judge and improve item quality.

PURPOSES

Find each item's difficulty level

Find each item's discriminating power

Find each distracter's effectiveness

Decide: retain, modify, or reject

PROCEDURE — KELLEY'S UPPER/LOWER 27% METHOD

1

Score all try-out sheets; arrange in descending order of total score

2

Take the top 27% as the Upper Group (U) and bottom 27% as the Lower Group (L)

3

For each item, count correct responses in U and in L

4

Apply the Difficulty & Discrimination formulae (next slides)

4. Improving Quality Through Item Analysis

23

24 of 29

4.3

Item Difficulty Index (p)

p = (Rᵤ + Rₗ) / (Nᵤ + Nₗ) × 100

Rᵤ, Rₗ = number correct in Upper / Lower group.

Nᵤ, Nₗ = group sizes (if equal, p = (Rᵤ+Rₗ)/2N × 100).

Definition: the proportion (%) of examinees who answered the item correctly.

Difficulty (p)

Interpretation

Above 80%

Very easy item

60% – 80%

Easy item

40% – 60%

Average / moderate difficulty — IDEAL

20% – 40%

Difficult item

Below 20%

Very difficult item

Items near 0% or 100% barely differentiate examinees; p around 40–60% is most desirable for a norm-referenced test.

4. Improving Quality Through Item Analysis

24

25 of 29

4.4

Item Discrimination Index (D)

D = (Rᵤ − Rₗ) / N

Rᵤ, Rₗ = number correct in Upper / Lower group.

N = number of examinees in each group (Nᵤ = Nₗ = N).

Definition: the extent to which an item differentiates high-achievers from low-achievers.

D Value

Quality of the Item

0.40 and above

Excellent discriminating item — retain as is

0.30 – 0.39

Good item — retain

0.20 – 0.29

Fair / marginal item — needs revision

Below 0.20

Poor item — revise substantially or reject

Negative value

Defective item — reject or thoroughly review

A negative D is a serious warning sign — an ambiguous stem, wrong key, or misleading option may be confusing competent examinees.

4. Improving Quality Through Item Analysis

25

26 of 29

4.5

Item Distracter Analysis

Examines how each incorrect option (distracter) of an MCQ actually functioned — how many, and which type of, examinees chose it.

AN EFFECTIVE DISTRACTER SATISFIES BOTH CONDITIONS:

1. Chosen by at least a few examinees (≈ 5%+) — an option nobody picks is ‘non-functional’

2. Chosen by relatively more of the Lower group than the Upper group — a plausible misconception

COMMON DEFECTS REVEALED

A distracter chosen by nobody — irrelevant / implausible; replace it

A distracter chosen more by the Upper group — the ‘key’ may be debatable or the option partially correct

The correct answer chosen more by the Lower group — indicates a flawed key or misleading stem

4. Improving Quality Through Item Analysis

26

27 of 29

4.6

Using the Results — Retain, Revise or Reject

RETAIN

Difficulty ~30–70%, discrimination ≥0.30, and functioning distracters — include as is.

REVISE

Moderate difficulty/discrimination, or one non-functional distracter — modify wording, key or an option, then re-try.

REJECT

Very poor / negative discrimination, or extreme difficulty with no scope for revision — drop the item.

Systematic item analysis, applied cyclically across successive try-outs, progressively improves the reliability, validity and discriminating power of the final standardized test.

4. Improving Quality Through Item Analysis

27

28 of 29

SUMMARY

Key Takeaways

1

Test construction & standardization is a scientific, sequential process producing a valid, reliable, objective and practicable measuring instrument with established norms.

2

Essay, short-answer and objective items each trade off content coverage, scoring objectivity, and the level of thinking measured.

3

The Achievement Test sequence runs: Planning → Design & Blueprint → Item Writing → Try-out → Item Analysis → Final Draft → Standardization.

4

Item analysis — Difficulty Index, Discrimination Index and distracter analysis — is the key statistical tool for retaining, revising or rejecting items.

Unit III | Test Construction & Standardization

28

29 of 29

Thank You

Unit III | Test Construction & Standardization

M.A. Education • Educational Measurement and Evaluation