ABCDEFGHIJKLMNOPQRSTUVWXYZ
1
Question ID SectionQuestion TextHelper Text
2
[Automatically Answered] Information about the Label Label First Publish Date
3
[Automatically Answered] Information about the Label Label Last Updated
4
[Automatically Answered] Information about the Label Label Version
5
[Automatically Answered] Information about the Label % of Label Questions Completed
6
1MetadataAuthor(s) of this Label (Name, Email, Affiliation, Connection to Dataset)
This information will show up in the header of the published Label
7
2MetadataPeople Consulted (name, affiliation)
This includes subject matter experts, community advocates, academics, etc.
8
3MetadataDataset Title
9
4MetadataProvide a short summary of this dataset.
Summarize this dataset in a few sentences, including the purpose and contents.
10
5MetadataWho owns the dataset? (Name, Organization)List both organization/company and individual when available.
11
6MetadataWho created the dataset? (Name, Organization)List both organization/company and individual when available.
12
7MetadataWho currently maintains the dataset? (Name, Email, Affiliation)List both organization/company and individual when available.
13
8MetadataIndicate which version of the dataset is being describedIf the dataset has multiple versions, DOI, etc.
14
9MetadataIs the dataset pubicly available?
15
10MetadataWhat is the license under which the dataset is made available?If possible, link to license description.
16
11MetadataWhen was this dataset created/published?
17
12Just the FactsIndicate up to five keywords relating to this datasetThis could include geography, domain, or other characteristics
18
13Just the Facts
Does the dataset have a metadata repository or data dictionary? If yes, provide the link or access point and if not, explain the features in this dataset.
If yes, provide the link and if not, explain the features in this dataset.
19
14Just the FactsWhat is the format of the dataset?For example: csv, txt, etc.
20
15Just the FactsHow many instances are in the dataset?For example: 10,000 data points, 200 images, etc.
21
16Just the FactsOver what timeframe was the data collected?
For example, a dataset my be published in 2022 but collected between 2005-2010. If the data continues to be collected, choose "ongoing" for the end date.
22
17Just the FactsWhat was the process of data collection?
23
18Just the Facts
What tools, services, or technologies were used to collect the data? If there were multiple steps, please include all.
Tools and services would include: software API, hardware apparatus or sensor, app, door-to-door, etc. Also indicate how they were used.
24
19Just the FactsDoes the dataset include information from upstream sources?
Upstream sources can include: other datasets, research papers, tweets, websites, etc.
25
19aJust the Facts
Name these sources and provide access point to the upstream dataset(s) where possible. (Link, Description)
26
19bJust the Facts
Upload visuals, such as data flow diagrams, that would illustrate the relationship to upstream data.
27
20Just the FactsWill the dataset be updated?
For example, to correct labeling errors, add new instances, delete instances, etc.
28
21Just the FactsWho funds the collection of this data? (Name, Organization)
Include the names of companies, institutions, grant award numbers, etc.
29
22Just the FactsWho funds the management of this dataset? (Name, Organization)
Include the names of companies, institutions, grant award numbers, etc.
30
23Just the FactsHas the data been reviewed for technical quality?
For example: checking for data type consistency, gaps in data completion, etc.
31
23aJust the FactsProvide access points to any outputs from the technical review.
For example, quality reports, a list of quality metrics and expected distributions, etc.
32
24Just the FactsWere any ethical review processes conducted?
For example: institutional review board approval, internal or third party ethical review process, a report on the potential impact on data subjects, etc.
33
24aJust the Facts
Provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation.
34
25Just the FactsDoes this dataset include information about humans?
35
25aJust the FactsDoes the dataset identify any subpopulations?For example: age, gender, race or ethnicity, etc.
36
25bJust the FactsIndicate subpopulation distributions within the dataset where possible.
You may upload any visuals that indicate distributions if that is available as well.
37
25cJust the FactsDoes this contain individual-level data?
38
25dJust the FactsDoes this dataset contain confidential data?
39
25eJust the FactsWas consent given by subjects?
40
26Collected (Why) - Use CasesWithin what domain(s) is this intended to be used?Is it specific to a particular domain?
41
27Collected (Why) - Use CasesWhat was the original intended use for this dataset?
This could include inferences or conclusions from individual instances, as well as models created from trends and correlations in the dataset.
42
28Collected (Why) - Use Cases
Are there other uses, beyond those intended by the dataset producers, for which you could imagine this dataset being used repsonsibly?
This could include inferences or conclusions from individual instances, as well as models created from trends and correlations in the dataset.
43
29Collected (Why) - Use CasesWhere and for what purposes has this dataset been used?
Provide a description and links to papers or systems that use this dataset.
44
30Collected (Why) - Use CasesHow should this dataset not be used?
This could include inferences or conclusions from individual instances, as well as models created from trends and correlations in the dataset.
45
31Collected (Why) - Use CasesAre there domains or industries in which this dataset should not be used?
46
32Collected (Why) - Use CasesWhich communities, groups, or identities are represented in this dataset?
47
33Collected (Why) - Use CasesIs the dataset made available under certain restrctions?
48
33aCollected (Why) - Use CasesList restrictions hereInclude any legal, usage, or license retrictions here.
49
34Collected (Why) - Use Cases
What concerns might you have about extrapolating trends or making generalized inferences from this dataset at a population level?
Broader populations could include geographic populations, demographic populations, or other groupings
50
35Collected (Why) - Use Cases
Describe any known mitigation strategies, technical or otherwise, to address the concerns stated above.
Mitigation strategies could include data imputation or removal techniques, grouping approaches, or domain-specific knowledge.
51
36Collected (Why) - Use Cases
Are there concerns around using this data to make decisions or predictions at the individual level?
Aggregate data may not capture variation across individuals, and therefore decisions made with that data may lack important information. In some cases, this lack of information can make those decisions incomplete, inaccurate, or discriminatory.
52
37Collected (Why) - Use Cases
Describe any known mitigation strategies, technical or otherwise, to address the concerns stated above.
Mitigation strategies could include data imputation or removal techniques, grouping approaches, or domain-specific knowledge.
53
38Collected (What)
Describe any cultural or domain assumptions in data the field definitions that are not made explicit in the data dictionary.
This could include: representing information on a continuum in a set of categories, definitions that are domain-specific, grouping multiple categories into one broader category, assuming certain concepts are common knowledge across cultures/domains, etc.
54
39Collected (What)Are there any proxy characteristics in this dataset?
A proxy characteristic is a variable or field that is implicitly connected to another variable or field, often one that is sensitive or connected to sensitive features (e.g. protected categories such as race, gender, sex).
55
40Collected (What)
Did people from the communities, groups, or identities represented in the dataset participate in the process of defining and planning this dataset?
This could include the participation of the racial, gender, socioeconomic, and ability, as well as geographic, identities that are represented in the dataset.
56
41Collected (What)
Describe any domain-specific knowledge that is necessary to ensure the dataset is used as intended.
This could include knowledge about terms, processes, or strategies for cateogrization that are domain-specific.
57
42Collected (What)Is there any content in this dataset that might be sensitive or upsetting?
58
42aCollected (What)Describe the sensitive content.
59
43Collected (How)
Were there any protocols for collecting and labeling data that required interpretation by individuals or trained systems?
This could include categorizing or labeling data manually using a protocol or leveraging a trained system.
60
43aCollected (How) Describe the protocols.
61
44Collected (How)
Did people from the communities, groups or identities represented in the dataset participate in the data collection process?
This could include the participation of the racial, gender, socioeconomic, and ability, as well as geographic, identities that are represented in the dataset.
62
44aCollected (How) Elaborate on this if possible.
63
45Collected (How) Are there any other representation issues that might be reflected in the data?
In many cases a dataset is a sample that represents a larger population; it is thus important to indicate where undersampling, oversampling, etc. may have taken place.
64
46Processed (How)
Describe any protocols followed to impute or fill in missing data, and indicate which fields were affected and in what quantity (%).
65
47Processed (How)
Describe any protocols followed to manipulate or adjust existing data, and indicate which fields were affected and in what quantity (%).
This includes actions such as cleaning, transforming, grouping data.
66
48Processed (How) Describe any data that is missing and the reason, if known.
This includes both excluded items and lost data, such as unreadable data, assumed false survey answers, information retrieval malfunction, etc.
67
49Processed (How) Was the raw data saved in addition to the preprocessed data?
68
50Known Issues Provide any access points to documentation of issues related to this dataset.
Include the URL or API endpoint where the information can be accessed.
69
51Known Issues Any other concerns or issues with this dataset?
70
52
Upstream Dataset Information (repeated per dataset)
How familiar are you with the intended use of this dataset?
71
53
Upstream Dataset Information (repeated per dataset)
How familiar are you with way that this data was collected?
72
53a
Upstream Dataset Information (repeated per dataset)
Describe any known issues surrounding how this data was collected
This includes questions of: respresentation, sampling, labeling, communities engaged or excluded, any other domain interpretation or knowledge
73
54
Upstream Dataset Information (repeated per dataset)
How familiar are you with how this data was cleaned and processed?
74
55
Upstream Dataset Information (repeated per dataset)
Is there anything else that someone who is using this data should know about this upstream dataset?