1 of 86

Human-Computer Interaction

saadh.info/hci

Week 10 (Thursday): Hypothesis Testing - When to use a test?

1

2 of 86

Attendance and Agenda

  1. Hypothesis testing
    • Which test?
    • When?

2

3 of 86

Announcements

  • Please submit the teammate evaluation form.
  • Submit assignment 2.
  • Please go back and check if you finished all quizzes. There are some missing.
  • Start working on milestone 2!

3

4 of 86

Good Design, Bad Design Examples

4

5 of 86

3. Experimental Research in HCI

Error bars show

±1 standard deviation

6 of 86

Experimental Research in HCI

6

Scientific Foundation

Experiment Design

Hypothesis Testing

Demo and Assignment 3

7 of 86

General Rules

7

OK to compute....

Nominal

Ordinal

Interval

Ratio

frequency distribution

Yes

Yes

Yes

Yes

median and percentiles

No

Yes

Yes

Yes

addition or subtraction

No

No

Yes

Yes

mean or standard deviation

No

No

Yes

Yes

ratio, or coefficient of variation

No

No

No

Yes

8 of 86

How Many IVs?

  • An experiment must have at least one independent variable
  • Possible to have 2, 3, or more IVs
  • But the number of “effects” increases rapidly with the size of the experiment:

  • Advice: Keep it simple (1 or 2 IVs, 3 at the most)

8

9 of 86

Dependent Variable

  • A dependent variable is a measured human behaviour (related to an aspect of the interaction involving an independent variable)
  • “Dependent” because it depends on what the participant does
  • Examples:
    • task completion time, speed, accuracy, error rate, throughput, target re-entries, task retries, presses of backspace, etc.
  • Dependent variables must be clearly defined
    • Research must be reproducible!

9

10 of 86

Control Variable

  • A control variable is a circumstance (not under investigation) that is kept constant while testing the effect of an independent variable
  • More control means the experiment is less generalizable (i.e., less applicable to other people and other situations)
  • Research question: Is there an effect of font color or background color on reading comprehension?
    • Independent variables: font color, background color
    • Dependent variable: comprehension test scores
    • Control variables
      • Font size (e.g., 12 point)
      • Font family (e.g., Times)
      • Ambient lighting (e.g., fluorescent, fixed intensity)
      • Etc.

10

11 of 86

Random Variable

  • A random variable is a circumstance that is allowed to vary randomly
  • More variability is introduced in the measures (bad!), but the results are more generalizable (good!)
  • Research question: Does user stance affect performance while playing Guitar Hero?
    • Independent variable: stance (standing, sitting)
    • Dependent variable: score on songs
    • Random variables
      • Prior experience playing a real musical instrument
      • Prior experience playing Guitar Hero
      • Amount of coffee consumed prior to testing. Etc.

11

12 of 86

Confounding Variable

  • A confounding variable is a circumstance that varies systematically with an independent variable
  • Should be considered, lest the results are misleading
  • Research question: In an eye tracking application, is there an effect of “camera distance” on task completion time?
    • Independent variable: Camera distance (near, far)
      • Near camera (A): inexpensive camera mounted on eye glasses
      • Far camera (B): expensive camera mounted above system display
    • Dependent variable: task completion time
    • But, “camera” is a confounding variable: camera A for the near setup, camera B for the far setup
    • Are the effects due to camera distance or to some aspect of the different setups?

12

13 of 86

Within-subjects, Between-subjects

13

14 of 86

Within-subjects, Between-subjects

14

Within-subjects

Between-subjects

15 of 86

Latin Squares

15

2 x 2

4 x 4

3 x 3

5 x 5

16 of 86

Balanced Latin Square

  • With a balanced Latin square, each condition precedes and follows each other condition an equal number of times
  • Only possible for even-orders
  • Top row pattern: A, B, n, C, n – 1, D, n – 2, …

16

4 x 4

6 x 6

17 of 86

17

Letters Only

Keyboard

Letters + Word Prediction

Keyboard

= LO

18 of 86

18

18

= LO

19 of 86

Longitudinal Study – Results1

19

1 MacKenzie, I. S., Kober, H., Smith, D., Jones, T., & Skepner, E. (2001). LetterWise: Prefix-based disambiguation for mobile text entry. Proceedings of the ACM Symposium on User Interface Software and Technology - UIST 2001, 111-120, New York: ACM.

20 of 86

Cost-Benefit Trade-offs

  • New, improved techniques sometimes languish
  • Evidently, the benefit in the new technique is insufficient to overcome the cost in learning it (see below)

20

21 of 86

Hypothesis Testing

21

Sarah not being picked on the first day = ⅘ = 80%

Five weeks in a row = ⅘ * ⅘ * ⅘ * ⅘ * ⅘ = 0.32 = 32%

Twelve weeks in a row = (⅘)^12 = 0.044… 4.4% something is fishy

22 of 86

What is Hypothesis Testing?

  • … the use of statistical procedures to answer research questions
  • Typical research question (generic):

  • For hypothesis testing, research questions are statements:

  • This is the null hypothesis (assumption of “no difference”)
  • Statistical procedures seek to reject or accept the null hypothesis (details to follow)

22

23 of 86

Parametric vs. Non-parametric Tests

  • Four assumptions of parametric tests
    1. Does the data show a distribution, e.g., normal distribution?
    2. Is the variance between the two groups approximately equal?
    3. Is the data in each group randomly and independently sampled?
    4. Are there any outliers?

  • In general, less assumptions for non-parametric tests. Parametric tests are more powerful and sensitive, i.e., have better ability to distinguish between groups.

23

24 of 86

Statistical Procedures

  • Two types:
    • Parametric
      • Data are assumed to come from a distribution, such as the normal distribution, t-distribution, etc.
    • Non-parametric
      • Data are not assumed to come from a distribution
    • Lots of debate on assumptions testing and what to do if assumptions are not met (avoided here, for the most part)
    • A reasonable basis for deciding on the most appropriate test is to match the type of test with the measurement scale of the data (next slide)

24

25 of 86

Measurement Scales vs. Statistical Tests

  • Parametric tests most appropriate for…
    • Ratio data, interval data
  • Non-parametric tests most appropriate for…
    • Ordinal data, nominal data (although limited use for ratio and interval data)

25

26 of 86

Tests Presented Here

  • Parametric
    • Analysis of variance (ANOVA)
      • Used for ratio data and interval data
      • Most common statistical procedure in HCI research
  • Non-parametric
    • Chi-square test
      • Used for nominal data
    • Mann-Whitney U, Wilcoxon Signed-Rank, Kruskal-Wallis, and Friedman tests
      • Used for ordinal data

26

27 of 86

Analysis of Variance

  • The analysis of variance (ANOVA) is the most widely used statistical test for hypothesis testing in factorial experiments
  • Goal 🡪 determine if an independent variable has a significant effect on a dependent variable
  • Remember, an independent variable has at least two levels (test conditions)
  • Goal (put another way) 🡪 determine if the test conditions yield different outcomes on the dependent variable (e.g., one of the test conditions is faster/slower than the other)

27

28 of 86

Why Analyse the Variance?

  • Seems odd that we analyse the variance, but the research question is concerned with the overall means:

  • Let’s explain through two simple examples (next slide)

28

29 of 86

ANOVA Test OR F-test

  • Measures if an independent variable had a significant impact on the dependent variable
  • Gives us:
    1. F-statistic
    2. The degrees of freedom for F-statistic
    3. The associated p value

29

30 of 86

F-statistic Explained

30

10

12

18

24

36

11

14

19

23

38

12

13

17

25

37

Most of the differences are due to people and the drink did not make much of a difference

31 of 86

F-statistic Explained

31

29

29

30

31

31

17

18

19

19

20

10

11

12

12

13

Are the differences still due to people?

32 of 86

F-statistic Explained

  • Figure out how much of variance comes from
    • The variance between the groups
    • The variance within the groups

  • Calculate the ratio: F = between groups/within groups

  • The larger the ratio, the more likely it is that the groups have different means.

32

33 of 86

F-statistic Explained

33

29

30

31

31

29

28

29

27

30

29

25

28

29

27

29

The result of an ANOVA shows F(2, 12) = 4.27, p = 0.04

34 of 86

Degrees of Freedom

F(b,w) = …

b is the degrees of freedom for variance between group

w is the degree of freedom for variance within groups

b = number of groups - 1

w = total number of observations - number of groups

34

35 of 86

Null Hypothesis and Alternative Hypothesis

35

36 of 86

36

Example #1

Example #2

“Significant” implies that in all likelihood the difference observed is due to the test conditions (Method A vs. Method B).

“Not significant” implies that the difference observed is likely due to chance.

37 of 86

Example #1 - Details

37

Error bars show

±1 standard deviation

Note: SD is the square root of the variance

Note: Within-subjects design

38 of 86

Example #1 – ANOVA1

1 ANOVA table created by StatView (now marketed as JMP, a product of SAS; www.sas.com)

Probability of obtaining the observed data if the null hypothesis is true

Reported as…

F1,9 = 9.80, p < .05

Thresholds for “p”

    • .05
    • .01
    • .005
    • .001
    • .0005
    • .0001

39 of 86

How to Report an F-statistic

  • Notice in the parentheses
    • Uppercase for F
    • Lowercase for p
    • Italics for F and p
    • Space both sides of equal sign
    • Space after comma
    • Space on both sides of less-than sign
    • Degrees of freedom are subscript, plain, smaller font
    • Three significant figures for F statistic
    • No zero before the decimal point in the p statistic (except in Europe)

40 of 86

Example #2 - Details

Error bars show

±1 standard deviation

File: anova-ex2.txt

41 of 86

Example #2 – ANOVA

Reported as…

F1,9 = 0.626, ns

Probability of obtaining the observed data if the null hypothesis is true

Note: For non-significant effects, use “ns” if F < 1.0, or “p > .05” if F > 1.0.

42 of 86

Example #2 - Reporting

42

43 of 86

What if there are more conditions?

43

A

B

C

D

44 of 86

More Than Two Test Conditions

44

45 of 86

ANOVA

  • There was a significant effect of Test Condition on the dependent variable (F3,45 = 4.95, p < .005)
  • Degrees of freedom
    • If n is the number of test conditions and m is the number of participants, the degrees of freedom are…
    • Effect 🡪 (n – 1)
    • Residual 🡪 (n – 1)(m – 1)
    • Note: single-factor, within-subjects design

45

46 of 86

Post Hoc Comparisons Tests

  • A significant F-test means that at least one of the test conditions differed significantly from one other test condition
  • Does not indicate which test conditions differed significantly from one another
  • To determine which pairs differ significantly, a post hoc comparisons tests is used
  • Examples:
    • Fisher PLSD, Bonferroni/Dunn, Dunnett, Tukey/Kramer, Games/Howell, Student-Newman-Keuls, orthogonal contrasts, Scheffé
  • Scheffé test on next slide

46

47 of 86

Scheffé Post Hoc Comparisons

  • Test conditions A:C and B:C differ significantly (see chart three slides back)

47

48 of 86

Between-subjects Designs

  • Research question:
    • Do left-handed users and right-handed users differ in the time to complete an interaction task?
  • The independent variable (handedness) must be assigned between-subjects
  • Example data set 🡪

48

File: betweensubjects.txt

49 of 86

Summary Data and Chart

49

50 of 86

ANOVA

  • The difference was not statistically significant (F1,14 = 3.78, p > .05)
  • Degrees of freedom:
    • Effect 🡪 (n – 1)
    • Residual 🡪 (m – n)
    • Note: single-factor, between-subjects design

50

51 of 86

Two-way ANOVA

  • An experiment with two independent variables is a two-way design
  • ANOVA tests for
    • Two main effects + one interaction effect
  • Example
    • Independent variables
      • Device 🡪 D1, D2, D3 (e.g., mouse, stylus, touchpad)
      • Task 🡪 T1, T2 (e.g., point-select, drag-select)
    • Dependent variable
      • Task completion time (or something, this isn’t important here)
    • Both IVs assigned within-subjects
    • Participants: 12
    • Data set (next slide)

51

52 of 86

Data Set

52

53 of 86

Summary Data and Chart

53

54 of 86

ANOVA

54

Can you pull the relevant statistics from this chart and craft statements indicating the outcome of the ANOVA?

55 of 86

ANOVA - Reporting

55

56 of 86

Chi-square Test (Nominal Data)

  • A chi-square test is used to investigate relationships
  • Relationships between categorical, or nominal-scale, variables representing attributes of people, interaction techniques, systems, etc.
  • Data organized in a contingency table – cross tabulation containing counts (frequency data) for number of observations in each category
  • A chi-square test compares the observed values against expected values
  • Expected values assume “no difference”
  • Example research question:
    • Do males and females differ in their method of scrolling on desktop systems? (next slide)

56

57 of 86

Chi-square – Example #1

57

MW = mouse wheel

CD = clicking, dragging

KB = keyboard

58 of 86

Chi-square – Example #1

58

χ2 = 1.462

Significant if it exceeds critical value �(next slide)

59 of 86

Chi-square Critical Values

  • Decide in advance on alpha (typically .05)
  • Degrees of freedom
    • df = (r – 1)(c – 1) = (2 – 1)(3 – 1) = 2
    • r = number of rows, c = number of columns

59

χ2(2) = 1.462 (< 5.99 ∴not significant)

60 of 86

ChiSquareGUI Software

  • Embedded in GoStats
  • Available on HCI:ERP web site
  • Note: calculates p (assuming α = .05)

60

Demo

61 of 86

Chi-square – Example #2

  • Research question:
    • Do students, teachers, and parents differ in their responses to the question: Students should be allowed to use mobile phones during classroom lectures?
  • Data:

61

File: chisquare-ex2.txt

62 of 86

Chi-square – Example #2

  • Result: significant difference in responses (χ2 = 20.5, df = 2, p < .0001)
  • Post hoc comparisons reveal that opinions differ between students:parents (1:3) and teachers:parents (2:3)
  • No difference in opinions for students:teachers (1:2)
  • From GoStats… (demo in class)

62

1 = students, 2 = teachers, 3 = parents

63 of 86

Non-parametric Tests for Ordinal Data

  • Non-parametric tests are used most commonly on ordinal data (ranks)
  • Type of test depends on
    • Number of conditions 🡪 2 | 3+
    • Design 🡪 between-subjects | within-subjects

63

64 of 86

Non-parametric – Example #1

  • Research question:
    • Is there a difference in the political leaning of Mac users and PC users?
  • Method:
    • 10 Mac users and 10 PC users randomly selected and interviewed
    • Participants assessed on a 10-point linear scale for political leaning
      • 1 = very left
      • 10 = very right
  • Data (next slide)

64

65 of 86

Data (Example #1)

  • Means:
    • 3.7 (Mac users)
    • 4.5 (PC users)
  • Data suggest PC users more right-leaning, but is the difference statistically significant?
  • Data are ordinal (at least), ∴ a non-parametric test is used
  • Which test? (see below)

65

3.7 4.5

(column means)

66 of 86

Mann Whitney U Test1

66

Test statistic: U

Normalized z (calculated from U)

p (probability of the observed data, given the null hypothesis)

Corrected for ties

Conclusion:

The null hypothesis remains tenable: No difference in the political leaning of Mac users and PC users (U = 31.0, p > .05)

1 Output table created by StatView (now marketed as JMP, a product of SAS; www.sas.com)

67 of 86

MannWhitneyUGUI Software

  • Embedded in GoStats
  • Available on HCI:ERP web site

67

Demo

68 of 86

Non-parametric – Example #2

  • Research question:
    • Do two new designs for media players differ in “cool appeal” for young users?
  • Method:
    • 10 young tech-savvy participants recruited and given demos of the two media players (MPA, MPB)
    • Participants asked to rate the media players for “cool appeal” on a 10-point linear scale
      • 1 = not cool at all
      • 10 = really cool
  • Data (next slide)

68

69 of 86

Data (Example #2)

  • Means
    • 6.4 (MPA)
    • 3.7 (MPB)
  • Data suggest MPA has more “cool appeal”, but is the difference statistically significant?
  • Data are ordinal (at least), ∴ a non-parametric test is used
  • Which test? (see below)

69

6.4 3.7

(column means)

70 of 86

Wilcoxon Signed-Rank Test

70

Test statistic: Normalized z score

p (probability of the observed data, given the null hypothesis)

Conclusion:

The null hypothesis is rejected: Media player A has more “cool appeal” than media player B �(z = -2.254, df = 1, p < .05).

71 of 86

WilcoxonSignedRankGUI Software

  • Embedded in GoStats
  • Available on HCI:ERP web site

71

Demo

72 of 86

Non-parametric – Example #3

  • Research question:
    • Is age a factor in the acceptance of a new GPS device for automobiles?
  • Method
    • 8 participants recruited from each of three age categories: 20-29, 30-39, 40-49
    • Participants demo’d the new GPS device and then asked if they would consider purchasing it for personal use
    • They respond on a 10-point linear scale
      • 1 = definitely no
      • 10 = definitely yes

72

73 of 86

Data (Example #3)

  • Means
    • 7.1 (20-29)
    • 4.0 (30-39)
    • 2.9 (40-49)
  • Data suggest differences by age, but are differences statistically significant?
  • Data are ordinal (at least), ∴ a non-parametric is used
  • Which test? (see below)

73

7.1 4.0 2.9

(column means)

74 of 86

Kruskal-Wallis Test

74

Test statistic: H (follows chi-square distribution)

p (probability of the observed data, given the null hypothesis)

Conclusion:

The null hypothesis is rejected: There is an age difference in the acceptance of the new GPS device.�(χ2 = 9.605, df = 2, p < .01).

75 of 86

KruskalWallisGUI Software

  • Embedded in GoStats
  • Available on HCI:ERP web site

75

Demo

76 of 86

Post Hoc Comparisons

  • As with the analysis of variance, a significant result only indicates that at least one condition differs significantly from one other condition
  • To determine which pairs of conditions differ significantly, a post hoc comparisons test is used
  • From GoStats… (demo in class)

76

1 = 20-29 yrs, 2 = 30-39 yrs, 3 = 40-49 yrs

77 of 86

Non-parametric – Example #4

  • Research question:
    • Do four variations of a search engine interface (A, B, C, D) differ in “quality of results”?
  • Method
    • 8 participants recruited and demo’d the four interfaces
    • Participants do a series of search tasks on the four search interfaces (Note: counterbalancing is used, but this isn’t important here)
    • Quality of results for each search interface assessed on a linear scale from 1 to 100
      • 1 = very poor quality of results
      • 100 = very good quality of results

77

78 of 86

Data (Example #4)

  • Means
    • 71.0 (A), 68.1 (B), 60.9 (C), 69.8 (D)
  • Data suggest a difference in quality of results, but are the differences statistically significant?
  • Data are ordinal (at least), ∴ a non-parametric test is used
  • Which test? (see below)

78

71.0 68.1 60.9 69.8

(column means)

79 of 86

Friedman Test

79

Test statistic: H (follows chi-square distribution)

p (probability of the observed data, given the null hypothesis)

Conclusion:

The null hypothesis is rejected: There is a difference in the quality of results provided by the search interfaces (χ2 = 8.692, df = 3, p < .05).

80 of 86

FriedmanGUI Software

  • Embedded in GoStats
  • Available on HCI:ERP web site

80

Demo

81 of 86

Post Hoc Comparisons

  • Same as with Kruskal Wallis test
  • From GoStats… (demo in next class)

81

1 = interface A, 2 = interface B, 3 = interface C, interface D

82 of 86

Points of Discussion

  • Reporting the mean vs. median for scaled responses
  • Non-parametric tests for multi-factor experiments
  • Non-parametric tests for ratio-scale data

82

83 of 86

Controversy: Likert Data

  • Mean vs. Median?

  • Many HCI experiments use questionnaires with response items presented on a Likert Scale or some combination of numbers and verbal tags

  • Likert data is open to human interpretation

  • Is it interval data like temperature or ordinal data?

83

84 of 86

Controversy: Parametric vs. Non-parametric

  • Are time and errors in completing a task normally distributed?

  • Choices:
    1. Proceed with parametric test
    2. Transform of clean the data in some manner to correct the violation and then proceed with parametric test
    3. Use of a non-parametric test

84

85 of 86

Statistical Tests Cheat Sheet

85

Inherited from my advisor

Will be shared on Canvas

86 of 86

Attendance & Next Time

  • Hypothesis Testing

86