CSE 519: Data Science
Steven Skiena
Stony Brook University
Lecture 24: Human-centric Data Science
About the Final
Monday 12/16 from 11:15AM-1:45PM
Eng 145 + Frey 309 (overflow). Note your exam location:
SBU ID # <= 115999999 take it in Frey 309
SBU ID # >= 116000000 take it in Eng 145
About 12-15 short answer/quantitative questions.
Ideally drawn equally from topics in class / chapters in my book.
Best to study by doing problems in textbook.
Send me suggested problems if you want.
What are You Worried About AI?
Public/Ethical Concerns about AI
•Economic dislocation: will these machines take jobs from people?
•Privacy concerns: is too much of my private information being given to machines to build these models?
•Bias concerns: are these models learning the wrong thing from the training data?
•Agency issues: are people adequately in the decision loop to step in when models go astray?
Economic Dislocation
•Many blue collar jobs are threatened: taxi/truck drivers, security guards,…
•Many classes of white collar work are threatened: programmers, lawyers, translation/transcription, education, pathology/medicine…
•This double threat makes it hard to know what to train for!
•But: future jobs will involve working with AI systems to do better work. This suggests the need for a broad general education: communication, technology, and empathy.
•But: unemployment is about 4%, seemingly impossible 10 years ago.
•But: technology has always destroyed jobs, yet generated new and better jobs, in ways impossible to predict.
•The world will always change rapidly and you will have to adapt to it.
Scale Concerns
With increasing scale comes automation, which take people out of the decision process.
Who can be appealed to when machines make the same decision everywhere?
Privacy Concerns
Bias Concerns
Agency Issues
Weapons of Math Destruction
It is important to understand the potential harm data-driven models can cause.
Correlation is not causation, but models so-trained can trigger actions and feedback mechanisms, resulting in self-fulfilling prophecies.
It is important for us as data scientists to think about societal issues in a constructive way.
Ethics and AI
I think the definition gets right that ethics governs people’s behavior.
People build and apply technology, so the ethical demands are upon those who build and use AI systems.
Moral principles are often in tension with each other, and different people can reasonably hold to different beliefs and standards.
Properties of a Good Model
Teacher Ratings: A Terrible Model
Certain school systems fire teachers based on the test-scores of their students:
Societal / Ethical Issues in Big Data
Integrity in Communications and Modeling
Fight the temptation to inflate your results, because you know the model and audience:
Transparency and Ownership
Model-Driven Bias and Filters
Learning algorithms pick up biases from biased training data:
Maintaining Security of Big Data
There are ethical responsibilities to encrypt and delete data to avoid security breaches:
Maintaining Privacy in Aggregate
People are often identifiable even if their names, address, and ID number are deleted.
Data scientists must be careful with private data.
Aggregating Expertise
No person/program has all the answers.
Social media and other new technologies have made it easier to collect and aggregate opinions on a massive scale.
But how can we separate the wisdom of crowds from the cry of the rabble?
Francis Galton’s Ox
At a livestock fair in 1906, villagers were invited to guess the weight of a ox.
Galton observed that none of the almost 800 observers guessed the correct weight (1,178 pounds).
Yet the average guess was amazingly close: 1,179 pounds!
Wisdom of Crowds: Penny Demo
Wisdom of Crowds: Rules of Demo
How many pennies do I have in this jar?
Ten of you write your opinions on cards
Ten more tell me them one by one.
Who wants to bet?
Independent Guesses
537, 556, 600, 636, 1200, 1250, 2350, 3000, 5000, 11000, 15000
Median: 1250, Mean: 3739, Actual: 1879
The median is closer than any other guess.
Conditioned Guesses
A second group of students made guesses after seeing the first group’s answers:
750, 750, 1000, 1000, 1000, 1250, 1400, 1770, 1800, 3500, 4000, 5000
Median: 1325, Mean: 1935, Actual: 1879
Showing other people’s guesses substantially conditioned the distribution (no outliers)
Financial Stakes
Allowing people to bet on the outcome yielded two wagers:
1500, 2000 Actual: 1879
People willing to bet their own money on an event are by definition more confident in their selection.
When is the Crowd Wise?
Diversities of Opinion
Crowds only add information when there is disagreement.
A committee with perfectly correlated experts contributes nothing more than any one of them.
Be an incomparable element on the partial order of life.
Mechanisms for Aggregation
On numerical aggregation, the mean or median work if errors are symmetrically distributed.
But which is more robust: mean or median?
On classification problems, voting is the basic aggregation mechanism.
But should all votes be treated equally?
The Condorcet Jury Theorem
If the probability of each voter being correct is p>0.5, the probability of a majority of voters being correct P(n) is >p.
For p=0.51, a jury of 101 members is right 57%, P(1001)=0.73 and P(10001)=0.9999.
As n gets large enough P(n) approaches 1.
Arrow’s Impossibility Theorem
No election system for summing permutations of preferences as votes satisfies:
Crowdsourcing Services
Services like Amazon Turk, Prolific , and Appen(Crowdflower) provide the opportunity to hire large numbers of people for small amounts of piece work.
In life, you generally get what you pay for.
Getting people to do your bidding requires both incentives and clear instructions.
Good Uses of Turkers
Sheepmarket.com
What creative endeavors can you think of that people will do for $0.02 each?
Bad Uses of Turkers
Applications for Turkers? (Projects)
Gamification
Make things fun so that people work for free!
Motivational techniques include:
Games with a Purpose (GWAP)
Keys to success include: