1 of 22

API CAN CODE �Data in Learners’ Lives

Lesson 5: Evaluating Data Sources

This work was made possible through generous support from the National Science Foundation (Award # 2141655).

2 of 22

Warmup

Answer the following questions with your best guesstimates: �

How many Starbucks stores do you think there are in the US?�

How many in-person orders do you think Starbucks �received in 2022? �

How many mobile orders do you think Starbucks �received in 2022?�

How many items do you think Starbucks has on �their menu?

2

3 of 22

Lesson 1.4 Recap

  • We discussed primary and secondary sources of data�
  • We looked at examples of each, including the SelfieCity project which consolidates primary sources (selfie-takers) into a secondary source (the website with all the selfie data)�
  • We also considered data sources for our own local issues of interest!

3

4 of 22

The Coffee Dataset

  • My question: �Has coffee consumption increased over the last two decades?�
  • My dataset: �Starbucks yearly data from 2009 - 2022�
  • Note: this is a secondary source! It has been aggregated from a variety of primary sources!

4

5 of 22

Evaluating Data: The 5Vs for K-12

5

Velocity

The recency�of data curation

Variety

The types and structure of the data

Veracity

Accuracy, reliability, completeness, and bias of the data

Volume

The amount �of data available

Value

 The ability to extract meaningful insights

6 of 22

Evaluating Data: The 5Vs for K-12

Is your data source aligned with the 5Vs for K-12 Framework?

For each V, we will discuss it together and then evaluate the Starbucks dataset to see if it meets the standards of that V!

6

7 of 22

Evaluating Data: Volume

The amount of data available and whether it is sufficient to support the investigation at hand.

  • Datasets that are too small may result in incomplete or misleading conclusions, while very large datasets may be difficult to manage or interpret.

  • Think of insights or the different questions you �could draw/pose from datasets, looking at �different timespans (e.g., daily weather vs. �past 10 days vs. past 10 years)

7

Volume

The amount �of data available

8 of 22

Evaluating Data: Volume

  • How much data is included in the dataset? Or, How many data points/observations/rows are included?
  • Is there enough data to answer the research questions?
  • Is there enough data to draw meaningful insights or conclusions?
  • Is there too much data or too many �observations/rows? Or is some of it not �relevant or not related to the main question?

8

Evaluate the Starbucks dataset with respect to Volume.

9 of 22

Evaluating Data: Velocity

How current or recent the data is and whether it reflects a static snapshot or changes over time.

  • Some questions require up-to-date data (e.g., social media trends), whereas others can be addressed using historical data (e.g., rates of population change).

  • Think of a regional COVID-19 dataset that is �updated weekly vs. a regional dataset that is �updated daily. Which one is more appropriate �for identifying sudden spikes or long-term trends? �

9

Velocity

The recency�of data curation

10 of 22

Evaluating Data: Velocity

  • When was the data curated or last updated?
  • Does it includes real-time or recent data?
  • Is the dataset relevant to the investigation period?
  • How recent does the data need to be to be useful for the current investigation?

10

Evaluate the Starbucks dataset with respect to Velocity.

11 of 22

Evaluating Data: Variety

The types of data included (e.g., numerical, categorical, etc.) and how the data are structured and organized (e.g., tabular format, JSON, map)

  • Various types of data can support richer analysis, but can also be more complex to interpret.

  • Think of a mammal’s dataset that includes their �diet (categorical), speed (numerical), life span �(numerical), habitat description (text), and �latitude and longitude values (spatial). Which �variables are the most useful for investigation?

11

Variety

The types and structure of the data

12 of 22

Evaluating Data: Variety

  • What data types are included in the dataset (e.g., numerical, categorical, text, images, etc.)?
  • Is the dataset structured and organized (e.g., tabular format, JSON, map)?
  • Is the dataset well-documented? Are there details about the variables included?
  • Is the data format consistent across observations �and easy to interpret with the available tools?

12

Evaluate the Starbucks dataset with respect to Variety.

13 of 22

Evaluating Data: Veracity

The accuracy, reliability, completeness, and potential biases in the dataset.

  • Inaccurate, incomplete or biased data can lead to false conclusions.

  • Let’s break it down and further understand each �dimension separately.

13

Veracity

Accuracy, reliability, completeness, and bias of the data

14 of 22

Evaluating Data: Accuracy

  • What do we know about where the data came from? For example, how was the data collected? By whom? For what purpose(s)?

  • Is the dataset accurate compared to known fact or external source?

Accurate, Reliable

Accurate, Unreliable

Inaccurate, Reliable

Inaccurate, Unreliable

14

Evaluate the Starbucks dataset �with respect to Veracity: Accuracy.

15 of 22

Evaluating Data: Reliability

  • Is the data source reliable? Is the source consistently measured?�
  • Some measuring instruments and designs may reliably measure data; others may introduce measurement variability (such as a malfunctioning thermometer or poorly worded question on a survey)

Accurate, Reliable

Accurate, Unreliable

Inaccurate, Reliable

Inaccurate, Unreliable

15

Evaluate the Starbucks dataset �with respect to Veracity: Reliability.

16 of 22

Evaluating Data: Completeness

  • Is there any missing data?�
  • Check to make sure the dataset does not have a lot of missing data points – particularly if they are enough to make the sample small, or if they are systematically missing in certain populations.

16

Evaluate the Starbucks dataset �with respect to Veracity: Completeness

17 of 22

Evaluating Data: Data Biases

  • Is the data potentially biased?�
  • Are there potential biases in data sampling, reporting, or measurement? Are there any under- or over-represented populations? Are responses influenced by social �context? Are tools used systematically mismeasuring �a phenomenon?

  • If there are biases, how might these biases �potentially influence the results?

17

Evaluate the Starbucks dataset �with respect to Veracity: Data Biases

18 of 22

Evaluating Data: Value

The relevance and usefulness of the data in answering a given question or generating meaningful insights.

  • Even well-structured, accurate datasets may not be useful for answering a specific question and drawing meaningful insights.

  • A dataset is valuable if it helps generate �evidence in support of (or against) an �explanation.

18

Value

 The ability to extract meaningful insights

19 of 22

  • Does the dataset contain the necessary information to address the scientific question or the phenomenon under study?
  • Can it be analyzed to derive meaningful insights or evidence-based claims?
  • Which variables are most helpful or insightful?
  • What additional data could make the dataset more useful?

19

Evaluate the Starbucks dataset with respect to Value.

Evaluating Data: Value

20 of 22

Source Evaluation Activity

  • Choose one of the sources you identified in the previous lesson.�
  • Search the Internet for a dataset for this source. You could use one of the sources we’ve seen in class before, like Kaggle or OpenDataDC.

Evaluate this dataset using the 5Vs, discussing each category with your small group.

20

21 of 22

Exit Ticket

Match these 5 scenarios to the V (from the 5Vs) that they do the best job of representing:�

  • A researcher collects 10,000 responses to a survey on favorite color.�
  • A student uses an API to collect live-updated data from Spotify on top hits.�
  • A professor uses a verified government database to find population statistics for an upcoming presentation.�
  • A company uses not only survey data, but also video-recorded interviews and field-testing data to draw conclusions about a new product.�
  • The San Francisco Zoo decides to use a new research paper specifically about gorillas, instead of one written about chimpanzees, to determine what food to give to their gorillas.

21

22 of 22

Thanks!

apicancode@umd.edu

22

This work was made possible through generous support from the National Science Foundation (Award # 2141655).

API Can Code is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike

4.0 International (CC BY-NC-SA 4.0) License