1 of 33

Virtual Face to Face 2021

WCAG 3 Conformance

1

2 of 33

Etiquette

  • Use queue: q+
  • Use topics on queue:
    • "q+ to say [topic]"
    • “q+ to ask [topic]”
    • Helps chairs manage conversation
  • Keep your points short
  • Avoid metaphors and allegories
  • Be mindful of offensive language

W3C Code of Ethics and Professional Conduct

2

3 of 33

Setting the stage (1 out of 2)

  • The FPWD is a starting point, not the end point
  • Today’s focus is testing and scoring building up to conformance
    • We have a Lot of work to do on the draft guidelines that were in the FPWD
    • We can’t do that work until we make the decisions for today
  • We have comments from both in and outside the group on the draft guidelines in the FPWD
    • For today, please use success criteria, guidelines, outcomes, etc from WCAG 2.x and the FPWD as examples
    • We will dive into specific issues with the FPWD draft guidelines at a future time
    • We believe the conversations that we have today will inform us on coming up with solutions to those issues

3

4 of 33

Setting the stage (2 out of 2)

  • We will discuss scope last
    • For prior discussions, we assume that whatever scope is declared as:
      • Testing is conducted against scope
      • Conformance is claimed against a declared scope
      • Altering the boundaries of scope will not impact the decisions we make today
  • We have set aside the first 60-90 minutes of next week’s AG call to further discuss today’s content
  • AG and Silver merging status
    • AG will be focusing on 3.0, smaller group work on 2.3
    • We will continue to hold joint meetings with the goal of merging in the near future (date tbd)

4

5 of 33

Schedule

  • Session 1: Testing
    • 13:00-15:00 UTC
    • Chair: Alastair, Topic Lead: Jeanne
  • Session 2: Scoring
    • 17:00-19:00 UTC
    • Chair: Chuck, Topic Lead: Jeanne
  • Session 3: Conformance
    • 21:00-23:00 UTC
    • Chair: Shawn, Topic Lead: Rachael

5

6 of 33

Overarching themes

These themes span all our conversations

  • Need for simplicity and objectivity balanced with need for flexibility to address disparities between functional needs
  • Need to support regulations
    • Can’t fully address with this audience
    • Need large scale conversation with regulators and legal experts
    • The level that conformance is placed at can vary and assumptions affect conversation
  • Need to support both big and small entities

6

7 of 33

WCAG 3 structure

  • Guidelines
    • High-level, plain-language version of the content
    • Example: Use sections, headings, and sub-headings to organize your content.
  • Outcomes
    • Testable criteria and include information on how to score the outcome
    • Example: Convey hierarchy with semantic structure
  • Methods
    • Detailed information on how to meet the outcome, code samples, working examples, resources, etc
    • Example: Semantic headings (HTML)
  • Functional Needs
            • Describes a specific gap in one’s ability, or a specific mismatch between ability and the designed environment or context
            • Draft Functional Needs

7

8 of 33

Requirements review (1 out of 2)

4.1 Multiple ways to measure: All WCAG 3.0 guidance has tests or procedures so that the results can be verified. In addition to the current true/false success criteria, other ways of measuring (for example, rubrics, sliding scale, task-completion, user research with people with disabilities, and more) can be used where appropriate so that more needs of people with disabilities can be included.

4.2 Flexible maintenance and extensibility: Create a maintenance and extensibility model for guidelines that can better meet the needs of people with disabilities using emerging technologies and interactions. The process of developing the guidance includes experts in the technology.

4.3 Multiple ways to display: Make the guidelines available in different accessible and usable ways or formats so the guidance can be customized by and for different audiences.

4.4 Technology neutral: Guidance should be expressed in generic terms so that they may apply to more than one platform or technology. The intent of technology-neutral wording is to provide the opportunity to apply the core guidelines to current and emerging technology, even if specific technical advice doesn't yet exist.

8

9 of 33

Requirements review (2 out of 2)

4.5 Readability/Usability: The core guidelines are understandable by a non-technical audience. Text and presentation are usable and understandable through the use of plain language, structure, and design.

4.6 Regulatory Environment: The Guidelines provide broad support, including structure, methodology, and content that facilitates adoption into law, regulation, or policy, and clear intent and transparency as to purpose and goals, to assist when there are questions or controversy.

4.7 Motivation: The Guidelines motivate organizations to go beyond minimal accessibility requirements by providing a scoring system that rewards organizations which demonstrate a greater effort to improve accessibility.

4.8 Scope: The guidelines provide guidance for people and organizations that produce digital assets and technology of varying size and complexity. This includes large, dynamic, and complex websites. Our intent is to provide guidance for a diverse group of stakeholders including content creators, browsers, authoring tools, assistive technologies, and more.

9

10 of 33

Session 1: Testing

10

11 of 33

What types of tests to include?

  • Granular testing
    • Automated Testing (WCAG 2.x, FPWD)
    • Subjective but clearly defined tests (WCAG 2.x, FPWD)�
  • Holistic Testing
    • Heuristic testing (FPWD)
    • AT testing
    • People-in-seats testing
    • Maturity model

11

12 of 33

Which tests to include in conformance?

  • Pros and cons of each:
    • Automated Testing (WCAG 2.x, FPWD)
    • Subjective but clearly defined tests (WCAG 2.x, FPWD)
    • Heuristic testing (FPWD)
    • AT testing
    • People-in-seats testing
    • Maturity model
  • Next steps
    • Should we remove subjectivity from WCAG3?
    • What tests do we include in WCAG3?

12

13 of 33

Session 2: Scoring

13

14 of 33

Etiquette

  • Use queue: q+
  • Use topics on queue:
    • "q+ to say [topic]"
    • “q+ to ask [topic]”
    • Helps chairs manage conversation
  • Keep your points short
  • Avoid metaphors and allegories
  • Be mindful of offensive language

W3C Code of Ethics and Professional Conduct

14

15 of 33

Overarching themes

These themes span all our conversations

  • Need for simplicity and objectivity balanced with need for flexibility to address disparities between functional needs
  • Need to support regulations
    • Can’t fully address with this audience
    • Need large scale conversation with regulators and legal experts
    • The level that conformance is placed at can vary and assumptions affect conversation
  • Need to support both big and small entities

15

16 of 33

Conclusions from Session 1

  • Resolution: For WCAG 3, testing will aim to improve inter-tester reliability and will work on testing to measure this
  • Agreed that the framing of inter-tester reliability (or reliability for short) was a more productive way to approach testing than discussing subjectivity, since WCAG 2 has subjectivity. We want to improve, if possible.
  • Agreed that AGWG members would assist the Alt Text Subgroup to work on writing qualitative evaluation that can be rated with improved inter-tester reliability. This can include breaking it into more granular tests and outcomes.

16

17 of 33

WCAG 3 structure

  • Guidelines
    • High-level, plain-language version of the content
    • Example: Use sections, headings, and sub-headings to organize your content.
  • Outcomes
    • Testable criteria and include information on how to score the outcome
    • Example: Convey hierarchy with semantic structure
  • Methods
    • Detailed information on how to meet the outcome, code samples, working examples, resources, etc
    • Example: Semantic headings (HTML)
  • Functional Needs
            • Describes a specific gap in one’s ability, or a specific mismatch between ability and the designed environment or context
            • Draft Functional Needs

17

18 of 33

Comparison to WCAG 2.x Structure

18

WCAG 2

WCAG 3

Guidelines

Guidelines

Success criteria

Outcomes

Techniques

Methods

Understanding

How To

Principles

Tags (TBD)

19 of 33

Conformance Comparison - WCAG 2

19

WCAG 2

WCAG 3

Evaluate by page

Evaluate by site or product (or subset)

A, AA, AAA

Critical Errors

Perfection or fail

Point System

AA is mostly used for regulations

Bronze will be recommended for regulations

Success criteria have the same true/false evaluation

Guidelines are customized for the tests and scoring that is most appropriate.

20 of 33

Disability equity

  • How do we treat different disabilities equally when the needs are uneven?
  • FPWD:
    • Run tests and score at outcome level
    • Average scores at total and by functional need categories
    • Threshold for passing at both total and functional need
  • This group agreed that we would not weight last deep dive
  • If someone has a new alternative to what we currently have, please write it up for consideration by the whole group at a future meeting

20

21 of 33

Scoring options at outcome level

  • Binary
    • Work well for tests that results in a pass or fail condition will be assigned a 100% or 0%.
    • When possible, this is a clear option
    • Example: Visual contrast
  • Percentages
    • Work well for tests at the element level that can be consistently counted (number passed / total number of instances);
    • For example, doing a % against potential headings doesn’t work well because the number of potential headings will vary by tester.
    • Example: Alt text presence in image tags
  • Rating Scales
    • Work well for tests that apply to content without clear boundaries.
    • The tighter the definition of each rating and associated concepts, the better a rating scale will work
    • Example: Captions
  • Points
    • Work similarly to a rating scale but aggregate a bit differently.

21

22 of 33

Example: Text Alternatives in FPWD

Two types of tests:

  • Automated for the presence of alternative text - percentage of total images
  • Manual for whether the images in the process or task have appropriate alternative text. - percentage of total images + critical error

As an example, a result of the tests: 83% of images have appropriate alternative text with no critical errors.

22

23 of 33

Example: Outcome rating

23

Rating

Criteria

Rating 0

Less than 60% of all images have appropriate text alternatives OR there is a critical error in the process

Rating 1

60% - 69% of all images have appropriate text alternatives AND no critical errors in the process

Rating 2

70%-79% of all images have appropriate text alternatives AND no critical errors in the process

Rating 3

80%-94% of all images have appropriate text alternatives AND no critical errors in the process

Rating 4

95% to 100% of all images have appropriate text alternatives AND no critical errors in the process

24 of 33

How to handle scoring?

  • Issues:
    • Testing degrees of success is not efficient: 508
    • Need for more than pass/fail but also need better explanation: 463
    • Need to allow some small failures: 448

  • How should the scoring be done at the outcome level?
    • Binary (WCAG 2.x, FPWD)
    • Percentage (FPWD at testing level)
    • Rating scale (FPWD at testing and outcome (confusing))
    • Points

24

25 of 33

How to handle the spoon problem?

  • Small problems add up to a major failure
  • FPWD Critical errors: When aggregated within a view or across a process stop a user from using the view or completing the process (example: a large amount of confusing, ambiguous language).
    • Con: Presently applies at the outcome level, and does not aggregate across multiple guidelines and outcomes
  • Other options:
    • Handle by setting or lowering the threshold of passing for certain functional needs
    • Identify guidelines that have this concern and create a combined threshold
  • What are other options?

25

26 of 33

Session 3: Conformance

26

27 of 33

Etiquette

  • Use queue: q+
  • Use topics on queue:
    • "q+ to say [topic]"
    • “q+ to ask [topic]”
    • Helps chairs manage conversation
  • Keep your points short
  • Avoid metaphors and allegories
  • Be mindful of offensive language

W3C Code of Ethics and Professional Conduct

27

28 of 33

Current state: FPWD

  • Bronze:
    • Outcomes similar to A, AA, AAA SC
    • Minimum Conformance Level
  • Silver
    • Some holistic testing needed but not defined
    • Bronze must be met first
  • Gold
    • Some holistic testing needed but not defined
    • Bronze must be met first

28

29 of 33

Conformance levels

  • Should there be a level of conformance lower than Bronze, that provides less benefit, but is in some ways easier to achieve?
    • A level based on automatable testing?
    • Should we have a structure where a site or product must meet all automatable tests before doing manual or qualitative tests?
  • Should the maturity model be a level or a separate document?
  • Should user testing be a level or a separate document?
  • What should motivate incremental improvements to products and/or advancing the field?
    • Levels based by on test type passed
    • Levels based on total score

29

30 of 33

How to address conformance challenges

  • Small entities
  • Large scale content
  • Rapidly changing content
  • Third party content
  • All software has bugs

Use Cases from Silver Exploring Conformance Solutions subgroup are creating another measure of the Conformance model.

March Report - Use cases potentially in the WCAG3 FPWD structure

April Report - Use cases not yet addressed in WCAG3 FPWD structure

30

31 of 33

How to handle the schedule constraints

  • Desire to publish within the next 3 years while adequately addressing the complexity of the task
  • Should WCAG 3 be presented as multiple documents or filters?
  • What should we do first?
    • Existing SC
    • Web only
    • Maturity Model
    • New technologies

31

32 of 33

Defining scope

  • Scope in FPWD
    • Views - a single interaction with content.
    • Process - a complete activity the user performs, comprised of one or more views
  • Issues raised
    • We need to better clarify process if critical errors are defined as stopping a process

32

33 of 33

33

Flowchart from the decisions of the August 2020 Joint meeting (if needed to illustrate some aspect of what was decided)