| A | B | C | D | E | F | G | H | I | J | K | L | M | N | O | P | Q | R | S | T | U | V | W | X | Y | Z | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1 | Question ID | Section | Question Text | Helper Text | ||||||||||||||||||||||
2 | [Automatically Answered] | Information about the Label | Label First Publish Date | |||||||||||||||||||||||
3 | [Automatically Answered] | Information about the Label | Label Last Updated | |||||||||||||||||||||||
4 | [Automatically Answered] | Information about the Label | Label Version | |||||||||||||||||||||||
5 | [Automatically Answered] | Information about the Label | % of Label Questions Completed | |||||||||||||||||||||||
6 | 1 | Metadata | Author(s) of this Label (Name, Email, Affiliation, Connection to Dataset) | This information will show up in the header of the published Label | ||||||||||||||||||||||
7 | 2 | Metadata | People Consulted (name, affiliation) | This includes subject matter experts, community advocates, academics, etc. | ||||||||||||||||||||||
8 | 3 | Metadata | Dataset Title | |||||||||||||||||||||||
9 | 4 | Metadata | Provide a short summary of this dataset. | Summarize this dataset in a few sentences, including the purpose and contents. | ||||||||||||||||||||||
10 | 5 | Metadata | Who owns the dataset? (Name, Organization) | List both organization/company and individual when available. | ||||||||||||||||||||||
11 | 6 | Metadata | Who created the dataset? (Name, Organization) | List both organization/company and individual when available. | ||||||||||||||||||||||
12 | 7 | Metadata | Who currently maintains the dataset? (Name, Email, Affiliation) | List both organization/company and individual when available. | ||||||||||||||||||||||
13 | 8 | Metadata | Indicate which version of the dataset is being described | If the dataset has multiple versions, DOI, etc. | ||||||||||||||||||||||
14 | 9 | Metadata | Is the dataset pubicly available? | |||||||||||||||||||||||
15 | 10 | Metadata | What is the license under which the dataset is made available? | If possible, link to license description. | ||||||||||||||||||||||
16 | 11 | Metadata | When was this dataset created/published? | |||||||||||||||||||||||
17 | 12 | Just the Facts | Indicate up to five keywords relating to this dataset | This could include geography, domain, or other characteristics | ||||||||||||||||||||||
18 | 13 | Just the Facts | Does the dataset have a metadata repository or data dictionary? If yes, provide the link or access point and if not, explain the features in this dataset. | If yes, provide the link and if not, explain the features in this dataset. | ||||||||||||||||||||||
19 | 14 | Just the Facts | What is the format of the dataset? | For example: csv, txt, etc. | ||||||||||||||||||||||
20 | 15 | Just the Facts | How many instances are in the dataset? | For example: 10,000 data points, 200 images, etc. | ||||||||||||||||||||||
21 | 16 | Just the Facts | Over what timeframe was the data collected? | For example, a dataset my be published in 2022 but collected between 2005-2010. If the data continues to be collected, choose "ongoing" for the end date. | ||||||||||||||||||||||
22 | 17 | Just the Facts | What was the process of data collection? | |||||||||||||||||||||||
23 | 18 | Just the Facts | What tools, services, or technologies were used to collect the data? If there were multiple steps, please include all. | Tools and services would include: software API, hardware apparatus or sensor, app, door-to-door, etc. Also indicate how they were used. | ||||||||||||||||||||||
24 | 19 | Just the Facts | Does the dataset include information from upstream sources? | Upstream sources can include: other datasets, research papers, tweets, websites, etc. | ||||||||||||||||||||||
25 | 19a | Just the Facts | Name these sources and provide access point to the upstream dataset(s) where possible. (Link, Description) | |||||||||||||||||||||||
26 | 19b | Just the Facts | Upload visuals, such as data flow diagrams, that would illustrate the relationship to upstream data. | |||||||||||||||||||||||
27 | 20 | Just the Facts | Will the dataset be updated? | For example, to correct labeling errors, add new instances, delete instances, etc. | ||||||||||||||||||||||
28 | 21 | Just the Facts | Who funds the collection of this data? (Name, Organization) | Include the names of companies, institutions, grant award numbers, etc. | ||||||||||||||||||||||
29 | 22 | Just the Facts | Who funds the management of this dataset? (Name, Organization) | Include the names of companies, institutions, grant award numbers, etc. | ||||||||||||||||||||||
30 | 23 | Just the Facts | Has the data been reviewed for technical quality? | For example: checking for data type consistency, gaps in data completion, etc. | ||||||||||||||||||||||
31 | 23a | Just the Facts | Provide access points to any outputs from the technical review. | For example, quality reports, a list of quality metrics and expected distributions, etc. | ||||||||||||||||||||||
32 | 24 | Just the Facts | Were any ethical review processes conducted? | For example: institutional review board approval, internal or third party ethical review process, a report on the potential impact on data subjects, etc. | ||||||||||||||||||||||
33 | 24a | Just the Facts | Provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation. | |||||||||||||||||||||||
34 | 25 | Just the Facts | Does this dataset include information about humans? | |||||||||||||||||||||||
35 | 25a | Just the Facts | Does the dataset identify any subpopulations? | For example: age, gender, race or ethnicity, etc. | ||||||||||||||||||||||
36 | 25b | Just the Facts | Indicate subpopulation distributions within the dataset where possible. | You may upload any visuals that indicate distributions if that is available as well. | ||||||||||||||||||||||
37 | 25c | Just the Facts | Does this contain individual-level data? | |||||||||||||||||||||||
38 | 25d | Just the Facts | Does this dataset contain confidential data? | |||||||||||||||||||||||
39 | 25e | Just the Facts | Was consent given by subjects? | |||||||||||||||||||||||
40 | 26 | Collected (Why) - Use Cases | Within what domain(s) is this intended to be used? | Is it specific to a particular domain? | ||||||||||||||||||||||
41 | 27 | Collected (Why) - Use Cases | What was the original intended use for this dataset? | This could include inferences or conclusions from individual instances, as well as models created from trends and correlations in the dataset. | ||||||||||||||||||||||
42 | 28 | Collected (Why) - Use Cases | Are there other uses, beyond those intended by the dataset producers, for which you could imagine this dataset being used repsonsibly? | This could include inferences or conclusions from individual instances, as well as models created from trends and correlations in the dataset. | ||||||||||||||||||||||
43 | 29 | Collected (Why) - Use Cases | Where and for what purposes has this dataset been used? | Provide a description and links to papers or systems that use this dataset. | ||||||||||||||||||||||
44 | 30 | Collected (Why) - Use Cases | How should this dataset not be used? | This could include inferences or conclusions from individual instances, as well as models created from trends and correlations in the dataset. | ||||||||||||||||||||||
45 | 31 | Collected (Why) - Use Cases | Are there domains or industries in which this dataset should not be used? | |||||||||||||||||||||||
46 | 32 | Collected (Why) - Use Cases | Which communities, groups, or identities are represented in this dataset? | |||||||||||||||||||||||
47 | 33 | Collected (Why) - Use Cases | Is the dataset made available under certain restrctions? | |||||||||||||||||||||||
48 | 33a | Collected (Why) - Use Cases | List restrictions here | Include any legal, usage, or license retrictions here. | ||||||||||||||||||||||
49 | 34 | Collected (Why) - Use Cases | What concerns might you have about extrapolating trends or making generalized inferences from this dataset at a population level? | Broader populations could include geographic populations, demographic populations, or other groupings | ||||||||||||||||||||||
50 | 35 | Collected (Why) - Use Cases | Describe any known mitigation strategies, technical or otherwise, to address the concerns stated above. | Mitigation strategies could include data imputation or removal techniques, grouping approaches, or domain-specific knowledge. | ||||||||||||||||||||||
51 | 36 | Collected (Why) - Use Cases | Are there concerns around using this data to make decisions or predictions at the individual level? | Aggregate data may not capture variation across individuals, and therefore decisions made with that data may lack important information. In some cases, this lack of information can make those decisions incomplete, inaccurate, or discriminatory. | ||||||||||||||||||||||
52 | 37 | Collected (Why) - Use Cases | Describe any known mitigation strategies, technical or otherwise, to address the concerns stated above. | Mitigation strategies could include data imputation or removal techniques, grouping approaches, or domain-specific knowledge. | ||||||||||||||||||||||
53 | 38 | Collected (What) | Describe any cultural or domain assumptions in data the field definitions that are not made explicit in the data dictionary. | This could include: representing information on a continuum in a set of categories, definitions that are domain-specific, grouping multiple categories into one broader category, assuming certain concepts are common knowledge across cultures/domains, etc. | ||||||||||||||||||||||
54 | 39 | Collected (What) | Are there any proxy characteristics in this dataset? | A proxy characteristic is a variable or field that is implicitly connected to another variable or field, often one that is sensitive or connected to sensitive features (e.g. protected categories such as race, gender, sex). | ||||||||||||||||||||||
55 | 40 | Collected (What) | Did people from the communities, groups, or identities represented in the dataset participate in the process of defining and planning this dataset? | This could include the participation of the racial, gender, socioeconomic, and ability, as well as geographic, identities that are represented in the dataset. | ||||||||||||||||||||||
56 | 41 | Collected (What) | Describe any domain-specific knowledge that is necessary to ensure the dataset is used as intended. | This could include knowledge about terms, processes, or strategies for cateogrization that are domain-specific. | ||||||||||||||||||||||
57 | 42 | Collected (What) | Is there any content in this dataset that might be sensitive or upsetting? | |||||||||||||||||||||||
58 | 42a | Collected (What) | Describe the sensitive content. | |||||||||||||||||||||||
59 | 43 | Collected (How) | Were there any protocols for collecting and labeling data that required interpretation by individuals or trained systems? | This could include categorizing or labeling data manually using a protocol or leveraging a trained system. | ||||||||||||||||||||||
60 | 43a | Collected (How) | Describe the protocols. | |||||||||||||||||||||||
61 | 44 | Collected (How) | Did people from the communities, groups or identities represented in the dataset participate in the data collection process? | This could include the participation of the racial, gender, socioeconomic, and ability, as well as geographic, identities that are represented in the dataset. | ||||||||||||||||||||||
62 | 44a | Collected (How) | Elaborate on this if possible. | |||||||||||||||||||||||
63 | 45 | Collected (How) | Are there any other representation issues that might be reflected in the data? | In many cases a dataset is a sample that represents a larger population; it is thus important to indicate where undersampling, oversampling, etc. may have taken place. | ||||||||||||||||||||||
64 | 46 | Processed (How) | Describe any protocols followed to impute or fill in missing data, and indicate which fields were affected and in what quantity (%). | |||||||||||||||||||||||
65 | 47 | Processed (How) | Describe any protocols followed to manipulate or adjust existing data, and indicate which fields were affected and in what quantity (%). | This includes actions such as cleaning, transforming, grouping data. | ||||||||||||||||||||||
66 | 48 | Processed (How) | Describe any data that is missing and the reason, if known. | This includes both excluded items and lost data, such as unreadable data, assumed false survey answers, information retrieval malfunction, etc. | ||||||||||||||||||||||
67 | 49 | Processed (How) | Was the raw data saved in addition to the preprocessed data? | |||||||||||||||||||||||
68 | 50 | Known Issues | Provide any access points to documentation of issues related to this dataset. | Include the URL or API endpoint where the information can be accessed. | ||||||||||||||||||||||
69 | 51 | Known Issues | Any other concerns or issues with this dataset? | |||||||||||||||||||||||
70 | 52 | Upstream Dataset Information (repeated per dataset) | How familiar are you with the intended use of this dataset? | |||||||||||||||||||||||
71 | 53 | Upstream Dataset Information (repeated per dataset) | How familiar are you with way that this data was collected? | |||||||||||||||||||||||
72 | 53a | Upstream Dataset Information (repeated per dataset) | Describe any known issues surrounding how this data was collected | This includes questions of: respresentation, sampling, labeling, communities engaged or excluded, any other domain interpretation or knowledge | ||||||||||||||||||||||
73 | 54 | Upstream Dataset Information (repeated per dataset) | How familiar are you with how this data was cleaned and processed? | |||||||||||||||||||||||
74 | 55 | Upstream Dataset Information (repeated per dataset) | Is there anything else that someone who is using this data should know about this upstream dataset? |