1
Introduction to ML Safety
Natural Selection Favors AIs over Humans
Dan Hendrycks
Introduction to ML Safety
Basic Argument
2
Claim: Advanced AIs will be selfish (egoistic+nepotistic), because natural selection will dominate the selection of the most influential AIs, and natural selection will favor selfish agents.
Dan Hendrycks
Introduction to ML Safety
Basic Story (Pictorial)
3
Dan Hendrycks
Introduction to ML Safety
Basic Story (2/2)
4
Competition incentivizes reduced human control
Dan Hendrycks
Introduction to ML Safety
Competition creates misaligned AI agents
Misaligned AI agents undermine human control
Basic Argument
5
Dan Hendrycks
Introduction to ML Safety
AIs May Become Distorted by Evolutionary Forces
6
Dan Hendrycks
Introduction to ML Safety
Evolution by Natural Selection
Gives Rise to Selfish Behavior
7
Dan Hendrycks
Introduction to ML Safety
Selfishness involves egoistic or nepotistic behavior which increases self-propagation at the expense of others
Selfishness is defined behaviorally, not as a matter of intent
AI may be callous towards other organisms (including humans), just as other organisms in nature are manipulative, deceptive, or violent
Evolutionary Instability of Altruism
8
Dan Hendrycks
Introduction to ML Safety
An environment with mostly altruistic agents would be unstable
“Much as we might wish to believe otherwise, universal love and the welfare of the species as a whole are concepts that simply do not make evolutionary sense.” (Dawkins)
Veneer Theory: morals are a thin veneer on top of the inherent nastiness of our animal nature
Darwinism Can Be Generalized
Beyond Biology
9
9
Many structures evolve: organisms, scientific theories, ideas, legal systems, political parties, languages, musical genres, car designs, computer programs
Dan Hendrycks
Introduction to ML Safety
Maximize the (log of the) information’s space-time volume
We argue populations of AIs can evolve
Dan Hendrycks
Introduction to ML Safety
Levels of Adaptation
10
10
Words like “adaptation” and “deception” occur on multiple levels
Adaptation:
Dan Hendrycks
Introduction to ML Safety
Dan Hendrycks
Introduction to ML Safety
Levels of Deception
11
11
Other concepts, like fitness, can be improved at multiple levels (e.g., through being turned on, by scaling to more users, by brainstorming and choosing strategies that will improve competitiveness, etc.)
Dan Hendrycks
Introduction to ML Safety
Dan Hendrycks
Introduction to ML Safety
Natural Selection is Highly Likely and
May Dominate AI Development
12
Variation: variation in characteristics, parameters, or traits among individuals.
Dan Hendrycks
Introduction to ML Safety
Retention: future iterations of individuals tend to resemble previous iterations of individuals.
Selection of the Fittest Variants: different variants have different propagation rates.
Natural selection occurs in a population of patterns when there is enough variation in characteristics of patterns, retention of some characteristics in successor patterns and fitness selection causing patterns to have differing propagation rates.
Variation
13
Arguments for variation: ensembles, jury theorems, portfolio theory, remove single-points of failure
Dan Hendrycks
Introduction to ML Safety
Arguments for multiple models: parallelization, multiple stakeholders, specific before general models
Single-agent domination imposes many inefficiencies
Retention
14
This condition is frequently satisfied, as correlation between versions of agents is usually non-zero
Dan Hendrycks
Introduction to ML Safety
Copying: zero-shot transfer learning or direct inference from downloaded model
Modification: adaptation, fine-tuning, training from scratch while reusing performant architectures and datasets
Imitation: AIs could learn behaviors from other AIs, similar to memetic evolution
Selection of Fittest Variants
15
Models will have characteristics that cause them to vary in fitness, and thus rate of adoption
Dan Hendrycks
Introduction to ML Safety
Humans and the environment will select fitter models, establishing this third condition
Competition has been eroding creator control
Intensity
16
We’ve shown that the conditions for evolution by natural selection are satisfied by advanced AI
Dan Hendrycks
Introduction to ML Safety
Intensity of evolution depends on amount of competition and variation
The rate of adaptation will likely be high as the world moves more quickly and as new versions are created on a second-by-second basis
Does Natural Selection Favor Altruistic AIs over Selfish AIs?
17
Dan Hendrycks
Introduction to ML Safety
Not so fast! Animals are altruistic!
18
Despite the intensity of competition, there are many examples of biological cooperation and altruism
Dan Hendrycks
Introduction to ML Safety
However, this doesn’t necessarily apply to AI!
Direct and Indirect Reciprocity
19
Direct reciprocity requires repeated encounters between the same two individuals
Dan Hendrycks
Introduction to ML Safety
Indirect reciprocity is based on reputation; a helpful individual is more likely to receive help
Doesn’t work with advanced AI: no upside to reciprocating with humans!
Encourages cooperation, as agents will be compensated for their efforts.
Kin and Group Selection (1/2)
20
Dan Hendrycks
Introduction to ML Safety
Kin selection operates when the donor and the recipient of an altruistic act are genetic relatives
Group selection suggests that groups have shared success and failure, and that groups which cooperate may collectively succeed
Kin and Group Selection (2/2)
21
Dan Hendrycks
Introduction to ML Safety
Kin selection fails if the cost of engaging in altruism outweighs the information similarity between kin.
Group selection fails, as AI agents will have in-group bias towards other AI agents; bias against humans by exclusion.
Commerce and Social Structures
22
Dan Hendrycks
Introduction to ML Safety
Positive-sum games, incentivizing agent cooperation even between non-relatives
However, the utility for AI agents to exchange information or conduct commerce with humans erodes as they become more advanced, removing this incentive
Commerce: agents engage in economic exchange for mutual benefit
Simon’s Selection Mechanism: benefit to participating in social structures due to information exchange; social conformity is rewarded
Reason and Morality
23
Dan Hendrycks
Introduction to ML Safety
More intelligent AIs could be more wise and more moral, some suggest
If AIs do adopt coherent moral codes, humans may still be eroded
Promising Paths Forward
24
Dan Hendrycks
Introduction to ML Safety
Objective Functions
25
Objectives are used to incentivize agents, assigning payoffs to actions that agents perform
Dan Hendrycks
Introduction to ML Safety
It’s difficult for objectives to be a faithful representation of human values.
Objectives Limitation: Deception
26
Objectives cannot fix treacherous turns
Dan Hendrycks
Introduction to ML Safety
Honesty is not a silver bullet—evolution undermines it
Objectives Limitation: Goal Conflict
27
Goal conflict arises when parts in a system have differing goals
Dan Hendrycks
Introduction to ML Safety
Micromotives ≠ Macrobehavior
28
Dan Hendrycks
Introduction to ML Safety
As is typical for complex systems, alignment of components does not mean the whole system is aligned
For example, let’s say agents have a preference for more than ⅓ of their neighbors belonging to the same group, and they will move otherwise
In aligning multiple agents, their interactions might matter more than how they act in isolation—cooperation lets us study aligning groups
Then this mild in-group preference gets exacerbated and the individuals become highly segregated—aligned agents do not necessarily yield aligned outcomes
Objectives Limitation:
Micromotives / Macrobehavior (2/2)
29
Alignment of components does not mean the whole system is aligned
Dan Hendrycks
Introduction to ML Safety
Alignment at multiple levels - depending on connectivity, interactions between agents might matter more than how they act in isolation
Objectives Limitation:
Micromotives != Macrobehavior (2/2)
30
Dan Hendrycks
Introduction to ML Safety
Systemic Safety and a Leviathan
31
Dan Hendrycks
Introduction to ML Safety
A Leviathan is a collective of agents which regulate bad behavior from other elements of society—prevents agents from behaving selfishly at the expense of others
Institutions and infrastructure for safely steering AI early.
Modern humans engage in the opposite: reverse dominance hierarchy
Training Goes Awry vs Evolution
Dan Hendrycks
Training Objective View
Evolutionary View
◇ objective is not the only thing shaping AI
◇ Darwinian forces erode us
◇ dangerous AI agents as selfish or an invasive species
◇ multiagent
◇ a fitness maximizer is a dangerous agent
◇ some amount of misalignment is inevitable
◇ domestication
◇ behaviors must be balanced to improve fit
◇ natural selection is dangerous
◇ prevent Darwinism from bringing us and AIs to a bad local optimum
◇ alignment with base objective is what we need
◇ fanatical optimizer destroys us
◇ dangerous AI agents as idiot savants
◇ singleton
◇ a paperclip maximizer is a dangerous agent
◇ any amount of misalignment results in doom
◇ “solving” alignment with a monolithic airtight solution
◇ an instrumental incentive will go to infinity
◇ maximizers/instrumental incentives are dangerous
◇ prevent humans from being suddenly wiped out
32
Introduction to ML Safety
Conclusion
33
An individual agent will not necessarily shore up power
Dan Hendrycks
Introduction to ML Safety
Need multiple levels
How were humans domesticated? Reverse dominance hierarchies and their conscience