1 of 21

CounterFace: A Synthetic Face Dataset

for Fine-Grained Counterfactual Evaluation

of Face Recognition Systems

Guruprasad Viswanathan Ramesh, Ashish Hooda, Shimaa Ahmed,

Harrison J. Rosenberg, Ramya Korlakai Vinayak, Kassem Fawaz

WI-PI Lab, University of Wisconsin-Madison

Wi-Pi Lab

2 of 21

Face Recognition is everywhere

Increasingly used for identity decisions with real consequences.

Surveillance

Airports & Borders

Retail

Personal Devices

1

3 of 21

Face Recognition in Action

Face Image 1

Face Image 2

Face Recognition Model

Face Recognition Model

Embedding 2

Embedding 1

Similarity Metric

 

Same Identity

Different Identity

 

2

4 of 21

Face Recognition Failures

  • Demographic Bias in Face Recognition has been widely reported

There is a need for fine-grained diagnostics.

  • Factors like aging, lighting, pose, and facial hair have also shown to cause failure

CNN; March 29, 2026; North Dakota, USA

NYT; Aug 6, 2023; Michigan, USA

3

5 of 21

Counterfactual Evaluation

Addition of attribute “mustache”

gai

Source x

Edited gai(x)

  • Explainable evaluation paradigm to assess robustness, biases and algorithmic fairness

  • Uses counterfactual examples obtained by making small human-interpretable changes to an input

4

6 of 21

Need for Synthetic Counterfactual Examples

  • Real counterfactual examples are hard to obtain at scale.
  • Synthetic counterfactual examples using modern generative models is the alternate
  • However, synthesized counterfactuals are not always reliable, and they need to be verified

5

While prompting different editing techniques to add the attribute “facemask”

7 of 21

Limitations of Prior Work

  • Human-in-the-loop → costly, does not scale

  • Checked only the intended attribute — confounders (other unrelated attributes) were not considered

Human annotators

GAN Editor

Verifier checks

“is attribute ai added or removed?”

Source face x

Edited face gai(x)

6

Prior Work: [1] Balakrishnan et al. (ECCV 2020) [2] Liang et al. (ICCV 2023)

8 of 21

Our work: Automated Verifier + Controlling Confounders

Automated verifier

Verifier checks three requirements:

✓ Validity

✓ Correctness

✓ Specificity

  • A validated automated verifier allows to scale to more attributes and demographics

  • Also allows minimizing confounders while selecting counterfactuals

  • We generate CounterFace, spanning 20 attributes and 8 demographics

Source face x

Edited face gai(x)

GAN Editor

7

9 of 21

1. Generate “Source Faces” for all 8 demographics (150 identities per demographic, 6 seeds per identity)

East Asian Male

White Male

Indian Male

Black Male

East Asian Female

White Female

Indian Female

Black Female

Generating CounterFace: Generating Candidate Face Pairs

8

10 of 21

2. Modify source face by applying or removing the intended attribute

Generating CounterFace: Generating Candidate Face Pairs

Attributes

or

Source Face

Attribute Detector

Remove

Add

Add

Modified Faces

9

11 of 21

1. Check Counterfactual requirement: Validity

Generating CounterFace: Rejecting Non-Compliant Candidates

Modified Faces

Artifact

Detector

Artifact Not Detected

Passes

Validity

Fails

Validity

Artifact Detected

Reject Candidate

Artifact Not Detected

Passes Validity

10

12 of 21

2. Check Counterfactual requirements: Correctness and Specificity

Generating CounterFace: Rejecting Non-Compliant Candidates

Attribute Detector

Candidate Pair for

Candidate Pair for

,

,

,

Attributes

or

,

or

,

,

,

Attributes

Attribute Responses

Attribute

modified correctly

Other attributes change

Attribute

modified correctly

Other attributes do not change

Passes

Correctness

Fails

Specificity

Passes

Correctness

Passes

Specificity

Reject Candidate

Select Candidate

11

13 of 21

  • 20 attributes & 8 demographics (160 attribute–demographic combinations)

  • 11,821 face pairs (75 pairs for all but 7 combinations)

Dataset

# Attr

# Demo

Verification

Open-sourced

Transect (Balakrishnan et al.)

6

4

Human

No

CausalFaces (Liang et al.)

4

6

Human

Yes

CounterFace (ours)

20

8

Automated

Yes (research)

  • 900-respondent post-hoc user study on a subset:
    • 96.9% pairs retain identity
    • 84.04% pairs meet all 3 counterfactual requirements

CounterFace: Dataset Details

12

14 of 21

Open-Source Models

Commercial Systems

AdaFace (Resnet-100 on MS1MV2 dataset)

AWS Rekognition Faces API

MagFace (Resnet-100 on MS1MV2 dataset)

Face++ Compare Faces API

FaceNet (Inception trained on VGGFace dataset)

ArcFace (Resnet-34 trained on MS1MV2 dataset)

  • Metrics: FNMR at FMR=0.1% after tuning decision threshold for each model/system

Evaluation: Models & Metrics

  • We evaluate four open-source models and two commercial systems

13

15 of 21

AWS & AdaFace are the most robust across all models

Results: General Performance Trends

14

16 of 21

Older open-source models, FaceNet and ArcFace show highest FNMR

Results: General Performance Trends

15

17 of 21

Results: Attribute-wise Performance

All models are less robust to rare attributes and attributes that occlude facial features

Models are fairly robust to non-occluding and peripheral attributes

16

18 of 21

Results: Attribute-wise Performance

Top2 models are also less robust to occluding attributes

17

19 of 21

Case Study: East Asian Male + Thick Beard

Even the top 2 models, AWS and AdaFace are less performant for East Asian demographic when thick beard is applied.

18

20 of 21

Ablation: Importance of Minimizing Confounders

All models lose performance with the relaxed dataset

Facial hair and occlusion attributes cause biggest FNMR delta on average

These results highlight the need for the Specificity Criterion that helps in minimizing impact of confounding factors in edits

19

21 of 21

Paper

Code

Dataset available upon request

Thank You, Any Questions??

20

Conclusion