1 of 77

Introduction to

AI Safety, Ethics, and Society

Normative Ethics

1

Introduction to AI SES

2 of 77

2

Introduction to AI SES

Introduction

Ethics is the branch of philosophy concerning questions of right and wrong, good and bad.

Moral decisions reflect our values, beliefs, and moral principles.

Ethics seeks to provide a systematic framework for these decisions.

Throughout this chapter we examine:

  1. Moral considerations
  2. Common ethical frameworks
  3. Ethical concepts central to AI safety, ethics, and society

Ethics

3 of 77

Roadmap

  1. Why learn about ethics?

  1. Normative factors: wellbeing, constraints, impartiality, etc.

  1. Moral Theories:
    • Utilitarianism
    • Deontology
    • Virtue Ethics
    • Social Contract Theory

  1. Moral Uncertainty

3

Introduction to AI SES

4 of 77

4

Introduction to AI SES

Why Learn About Ethics?

To develop a foundation for understanding AI safety discourse. This is important because…

  1. Widespread use and development of AI systems is taking place throughout various domains of human life.

2. Ensuring that AI systems are aligned with human values reduces the chance of bad outcomes.

3. AI systems generate ethical questions that are unique to their

technological nature and capabilities.

5 of 77

5

Introduction to AI SES

Ethics and Relativism

It may be that no moral principles or judgements are absolutely or universally correct; instead, they are only correct within a culture.

Moral Relativism is the view that there is no objective standard of morality that applies to all people across all situations.

Problems with moral relativism:

  1. Extremely harmful actions are do not permissible just because they are culturally relative.
  2. There may exist moral disagreements within cultures.

6 of 77

6

Introduction to AI SES

Ethics and Religion

Divine command theory is the view that the moral value of an action is determined solely by God’s commands because God is ‘good’.

But, if one should follow God’s word because God is good, then there must exist some moral qualities that are independent of God’s rules, for instance, ‘goodness’.

Therefore, divine command theory is false, and we cannot equate religion with morality.

7 of 77

Roadmap

  1. Why learn about ethics?

  1. Normative factors: wellbeing, constraints, impartiality, etc.

  1. Moral Theories:
    • Utilitarianism
    • Deontology
    • Virtue Ethics
    • Social Contract Theory

  1. Moral Uncertainty

7

Introduction to AI SES

8 of 77

Decomposing Goodness

8

Introduction to AI SES

Moral decisions often consider the intrinsic and instrumental goods that are at stake.

Instrumental goods are valuable because of the benefits or outcomes they provide – for example, money, power, career opportunities.

Intrinsic goods are valuable in and of themselves; common supposed examples are love, pleasure, beauty, and truth.

Some philosophers think there are many intrinsic goods, others accept none, or only one — wellbeing.

9 of 77

Wellbeing

9

Introduction to AI SES

How wellbeing is defined is debatable; however, it is commonly considered to be an intrinsic good.

Wellbeing generally refers to how well a person’s life is going for them, or whether a person is happy, healthy, and fulfilled.

Throughout this section, we discuss three common accounts of wellbeing:

  1. Wellbeing as pleasure.
  2. Wellbeing as preference satisfaction.
  3. Wellbeing as a set of objective goods.

10 of 77

Wellbeing as Pleasure

Wellbeing as pleasure: the balance of pleasure (or happiness) over pain (or suffering) is what constitutes wellbeing. Also called hedonism.

While we may all have different preferences and desires, pleasure seems to be universally valued. All other goods might be instrumental.

However, some philosophers argue against this view because it does not consider other ostensibly relevant factors to our wellbeing such as the pursuit of knowledge.

10

Introduction to AI SES

11 of 77

Wellbeing as Preference Satisfaction (1/2)

11

Introduction to AI SES

Wellbeing as preference satisfaction: what really matters for wellbeing is that our desires or preferences are satisfied.

Different types of preferences might differ greatly in different situations. Consider Alice’s preferences over different candidates in an election:

  1. Stated preference → “I prefer AI’s output X”
  2. Revealed preference → I chose AI’s output X
  3. Informed preference → I would have chosen AI’s output X if I knew…

Which preferences matter?

12 of 77

Wellbeing as Preference Satisfaction (2/2)

12

Introduction to AI SES

Stated preferences are outwardly expressed, but sometimes, they conflict with revealed preferences. A revealed preference illustrates our actions, not our intentions and they can be more useful than stated preferences.

A voter may express a preference for Google over ChatGPT, but when they actually have queries, they may more frequently use ChatGPT.

An informed preference considers all relevant information and can change as more information is acquired. These help us to uncover the reasoning behind stated/revealed preferences to better understand how they might change or conflict.

13 of 77

Wellbeing as Objective Goods

13

Introduction to AI SES

Objective goods are those things which are good independent of personal beliefs or opinions – they are universal.

An example set: moral goodness, rational activity, development of abilities, having children and being a good parent, knowledge, awareness of true beauty.

We may not always agree on what constitutes an objective good, and whether it can be considered universal.

14 of 77

Wellbeing and AI

14

Introduction to AI SES

AI influence on human wellbeing: AI chatbot companions could influence our wellbeing

  • If they aim to maximize revealed preferences, they may addict us and leave us unhappy
  • If they aim to maximize pleasure, they may give us amusing but shallow content
  • If they aim to maximize objective goods, they may push us to pursue self-development goals that we may not wish

Wellbeing of AIs: if AIs are conscious, then

  • Many coherently-acting AIs could have wellbeing according to a preference view of wellbeing
  • AIs that can reason, gain knowledge, or develop their abilities could have wellbeing according to an objective goods view

15 of 77

Wellbeing Summary

15

Introduction to AI SES

In this section, we have covered the following accounts of wellbeing:

  1. Wellbeing as pleasure
  2. Wellbeing as preference satisfaction
  3. Wellbeing as a set of objective goods

Though these views differ from one another they all share a common goal: promoting actions that result in morally good outcomes.

A morally good outcome will almost always aim to affect wellbeing to some degree.

16 of 77

Obligations, Constraints and Rights (1/2)

16

Introduction to AI SES

Constraints are actions that we’re morally prohibited from taking; common examples are killing or lying.

Obligations are actions that we’re morally required to take; common examples are keeping our promises or telling the truth.

Rights are claims that individuals may have over their community.

  • Rights that require certain actions from others are called “positive rights” → access to education, food, or shelter.
  • Rights that require others to abstain from certain behaviors are called “negative rights” → censorship or discrimination.

Some say AIs should eventually get negative rights but not positive rights.

17 of 77

Obligations, Constraints and Rights (2/2)

17

Introduction to AI SES

Obligations and constraints are often derived from respect for people’s rights. Individuals with rights could include:

  • Non-human individuals, such as AI and animals
  • Individuals that can experience pleasure and suffering
  • Individuals that have ‘interests’

Rights can be absolute and universal, or particular human rights vs. the right to vote in national elections.

18 of 77

Partiality and Impartiality (1/2)

Partial moral theories imply that we should value individual lives preferentially. These emphasize special obligations.

  • We have special obligations towards ourselves, our friends, our communities, or our families that we do not have toward strangers.

18

Introduction to AI SES

Impartial moral theories imply that we should value individual lives equally. These disregard special obligations.

Most modern moral theories are impartial, and almost every moral theory requires impartiality in at least some contexts.

19 of 77

Partiality and Impartiality (2/2)

Most modern theories agree that all human beings, regardless of their differences, are included in the moral circle – they are subjects of moral concern.

Some philosophers, such as hedonists, have also argued that non-human animals should be included, due to a capacity for suffering.

But perhaps we should extend the moral circle even further, considering entities that don’t have the same kinds of conscious experience as current humans, such as future humans or AIs.

Impartial AIs care for humans initially but come to care for them less.

19

Introduction to AI SES

20 of 77

20

Introduction to AI SES

Obligatory vs. Non-Obligatory Actions

Obligatory actions are those that we are morally obligated or required to perform – we have a moral duty to carry out these actions.

These actions can take the form of helping someone in distress or respecting human rights.

Non-obligatory actions are those that are not morally required or necessary.

These actions can still be morally good, for instance volunteering or donating to charity.

21 of 77

21

Introduction to AI SES

Permissible vs. Impermissible Actions

Impermissible actions are those that violate moral laws and are considered wrong (e.g., stealing or intentionally harming someone).

Permissible actions are those that are not impermissible – they can be divided into two categories: supererogatory or neutral.

  • Supererogatory actions go ‘above and beyond’ what’s morally required → donating to charity.
  • Neutral actions are morally inconsequential → taking a walk.

22 of 77

22

Introduction to AI SES

Praiseworthiness vs. Blameworthiness

Moral judgments regarding praiseworthiness and blameworthiness often consider moral responsibility and accountability.

Example: A toddler kicks their classmate → We do not blame the child if they were never taught otherwise.

Example: A billionaire donates to charity solely for tax benefits → We do not praise the billionaire because the reasons for donation are morally objectionable.

Such moral judgments do not imply anything about what is morally wrong, permissible, obligatory, or supererogatory.

23 of 77

Recap

Understanding ethics is important because AI systems—likely to have serious impacts on the world—must be used ethically. In particular, we must ensure that AI systems are aligned with human values.

Three conceptions of wellbeing suggest different ways to use AIs.

Moral considerations like impartiality, obligations, constraints, and rights give us a shared language to describe and debate the ethics of AIs and their impacts. These considerations are all relevant to common-sense morality and the following following moral theories.

23

Introduction to AI SES

24 of 77

Introduction to

AI Safety, Ethics, and Society

Normative Ethics

Part 2: Utilitarianism and Deontology

24

Introduction to AI SES

25 of 77

Roadmap

  1. Why learn about ethics?

  1. Moral considerations: wellbeing, obligations, partiality, et cetera.

  1. Moral Theories:
    • Utilitarianism
    • Deontology
    • Virtue Ethics
    • Social Contract Theory

  1. Moral Uncertainty

25

Introduction to AI SES

26 of 77

From Considerations to Theories

26

Introduction to AI SES

Moral considerations are part of metaethics: the underlying concepts and assumptions that make moral reasoning possible.

Normative ethics is about concrete moral standards and principles that govern how we ought to behave. Here, we consider the four most common theories in academic ethics:

  1. Utilitarianism
  2. Kantian Ethics/Deontology
  3. Virtue Ethics
  4. Social Contract Theory

27 of 77

Utilitarianism Overview

Utilitarianism seeks to maximize overall wellbeing → the right action is the one that increases overall wellbeing the most.

According to utilitarianism, the right action in any situation is the one which will increase overall wellbeing the most—not just for the people directly involved in the situation, but globally.

When we maximize wellbeing, we can make moral questions into empirical ones.

27

Introduction to AI SES

28 of 77

Utilitarianism: Drunk Driving (1/2)

28

Introduction to AI SES

Drunk driving. Amanda has had a few drinks, and is deciding whether to drive or take the bus home. What should she do?

Utilitarians would analyze the situation by…

  1. Listing all the possible outcomes of driving vs. not driving.
  2. Calculating the probabilities of each outcome.
  3. Calculating the ‘utility’ – wellbeing – brought about by each outcome.

Expected utility of each choice is the sum of the probabilities of each outcome multiplied by its utility.

29 of 77

Utilitarianism: Drunk Driving (2/2)

29

Introduction to AI SES

*the probability and utility values we assign here are arbitrary, and only serve to illustrate what such a calculation might look like.

Expected utility:

Amanda gets the bus → 1 x - 1 = - 1

Amanda drives home

0.95 x 1 + (0.05 x - 1000) = - 49.5

Amanda should take the bus!

Amanda’s action

Possible outcome(s)

Probability of each outcome

Utility

Amanda takes the bus.

Amanda is frustrated, the bus is slow, and she has to wait in the cold.

1

-1

Amanda drives home.

Amanda gets home safely, far sooner than she would have on the bus.

.95

+1

Amanda gets into an accident and someone is fatally injured.

.5

-1000

30 of 77

Utilitarianism’s Claims

Utilitarianism is a form of consequentialism → consequences determine whether an action is good or bad.

We measure* the effects of actions on wellbeing alone → wellbeing is the only intrinsic good.

All people have the same intrinsic moral worth → everyone’s wellbeing should be weighed impartially.

It is insufficient to do good, we must do what is best → we should maximize wellbeing.

*We measure wellbeing according to our choice of theory, e.g. Hedonism vs. Preference Satisfaction.

30

Introduction to AI SES

31 of 77

Critiques of Utilitarianism (1/3)

Utilitarianism is too demanding → If our money is more helpful to others than it is to us, there isn’t a utilitarian reason to keep it. For instance, Peter Singer has argued that we should donate at least a third of our income to charity.

Defences include:

  1. Utilitarianism is not as demanding in practice as it is in theory. Often, acting in a utilitarian way is quite “normal”, avoiding issues like burnout.
  2. We live in a demanding world. Maybe most of us should donate more.

Utilitarianism over-emphasizes wellbeingRobert Nozick famously highlights this point in his popular thought experiment.

31

Introduction to AI SES

32 of 77

Critiques of Utilitarianism (2/3)

32

Introduction to AI SES

Experience Machine: Consider an experience machine offering any desired sensation. Neuropsychologists could make you believe you're achieving great feats while you float in a tank, your brain connected to electrodes. Would you choose to live connected to this machine, pre-set with life's experiences? Most people say no.

You ‘wake up’ in an Experience Machine tomorrow. You can either forget this, returning to your current (simulated) life or return to your “real” one.

What if we reverse the thought experiment?

Most people choose to stay in the machine. This is status quo bias.

33 of 77

Critiques of Utilitarianism (3/3)

Utilitarianism requires intractable reasoning → we cannot reliably and precisely calculate future utility values. It is also impractical to expect this. The drunk driving example is idealized; for instance, we did not list all possible outcomes or consider long-term consequences.

33

Introduction to AI SES

The utilitarian response: a theory’s criterion of rightness doesn’t need to take the same form as its decision procedure.

A criterion of rightness → whatever a theory claims makes an action right.

A decision procedure → the process a theory recommends that individuals use to make decisions.

34 of 77

Deontology

34

Introduction to AI SES

Deontology’s key ideas:

  1. Focus on constraints and obligations, not the best outcome.
  2. Formulation of rules or principles that protect human autonomy.
  3. The rights and intentions of individuals are morally relevant.

Deontological rules from the Ten Commandments: “Thou shalt not kill”, “Thou shalt not steal”, “Honor thy mother and father”, etc.

Deontological theories are systems of moral rules, rights, duties and obligations which constrain our behaviour. These theories typically present themselves as improvements over consequentialism.

35 of 77

Deontology’s Critique of Utilitarianism

35

Introduction to AI SES

Consequentialism doesn’t allow options important life decisions like marriage or career choice will have “correct” utility-maximizing answers.

Deontologists say this view is too demanding – to preserve human autonomy, we provide a list of impermissible actions, and allow choice to govern the rest.

Consequentialism doesn’t forbid extreme actions → killing innocents, torturing – these actions aren’t always wrong for consequentialists.

Deontologists argue some actions are simply forbidden by appeal to universal constraints or obligations.

36 of 77

Deontology: The Doctrine of Double Effect

36

Introduction to AI SES

Deontology also emphasizes the intentions of agents in their moral considerations. Consider The Doctrine of Double Effect:

An agent is morally allowed to carry out actions that predictably lead to bad outcomes as long as they intend the good effect, but not the bad effect of the action.

Siamese twins: a pair of siamese twins needs a medical procedure, without which both of them would die. However, one of them will die as a consequence. The doctrine of double effect says the procedure is justified insofar as we intend to save one of the twins.

37 of 77

Deontology: Act/Omission

37

Introduction to AI SES

Intuitively, we do not hold a person responsible for what they did not do. In deontology, this is the Act vs. Omission distinction.

Stealing: While walking past a bank at night, Bob notices the night deposit box is open. Inside it, he sees a bag, filled with money. He decides to steal the bag. Alice sees Bob taking the money, but never bothers to report it.

Deontologists would not hold Alice responsible for Bob’s actions.

38 of 77

Critiques of Deontology

Deontology responds unconvincingly to moral catastrophes.

Nuclear Terrorism. The only way to prevent a nuclear terror attack is by torturing the perpetrator. But torture is impermissible to deontologists, regardless of the consequences.

In response, some some deontologists have adopted a threshold: when enough lives are at stake, the theory defers to consequentialism. However, there is no clear, non-arbitrary way to determine “enough”.

Deontology holds constraints like promises over vastly better outcomes.

38

Introduction to AI SES

39 of 77

Kant’s Ethics: The Categorical Imperative

39

Introduction to AI SES

Deontology can be traced back to Immanuel Kant, the 18th century Prussian philosopher. Kant’s ethical theory is called the Categorical Imperative. Under this theory…

  • Morality must be universal → each individual can rationally derive moral laws on their own.
  • All moral rules are categorical → they aren’t aiming at a goal, they are “In situation Y, I will do X” not “I will do X to get Z”.

He formulates the imperative in four ways. We will outline the Universal Law and Humanity formulations.

Immanuel Kant

40 of 77

Kant’s Universal Law Formulation (1/2)

40

Introduction to AI SES

To figure out if something we want to will is permissible, it must pass four stages of the universal law test.

Stage 1: Turn your proposed action into a rule

Stage 2: Turn the rule into one that applies to everyone

Stage 3: Is it possible to imagine a world where everyone follows the rule? Or is there a contradiction in conceiving of this world?

Stage 4: If there is no contradiction, would anyone will this rule?

In other words, we ask ourselves, “what if everyone did that?”, “what would the world look like?”, and “why should we live like this?”

41 of 77

Kant’s Universal Law Formulation (2/2)

41

Introduction to AI SES

Can the following rules be formulated as universal law?

I will not keep promises when doing so would inconvenience me

No → It fails at stage 3. This rule requires that the institution of promise-keeping exists. If everyone adopted this rule, promise-keeping would cease to exist; the rule is contradictory.

I will never lie under any circumstances

Yes → if everyone adopted this rule, the institution of honesty would continue to exist. We would will this rule because it preserves our autonomy: when we lie, we treat people as means to an end.

42 of 77

Kant’s Humanity Formulation

42

Introduction to AI SES

To have ‘humanity’ means being able to engage in autonomous, rational behaviour, and to choose your own projects.

The humanity formulation requires treating humanity as an end, not a means, emphasising the preservation of other’s autonomy.

The Urgent Lift. Alice wants a lift to go shopping. She lies to Bob, a stranger, telling him that her brother is having an allergic attack in town. She has his EpiPen and if he doesn’t give her a lift, her brother might die.

Alice uses Bob as a means to an end. Bob might have something more important to do than Alice’s shopping, but less critical than saving a life. The humanity formulation ensures that Bob can make this decision himself.

43 of 77

Critiques of Kant

43

Introduction to AI SES

Kant’s ethics are too extreme and ambiguous.

Mad axeman. Bob hears a knock on his door late at night. When he opens the door, a man with a wild look in his eye, and a bloody axe in his hand asks, “Is Alice in?”. Bob knows that Alice is sleeping upstairs, should he tell the man?

Kant would say that we should never lie to the mad axeman. Thus, his ethics can:

  • Fail to grasp the moral complexity of real-world situations.
  • Lead us to extreme moral conclusions.

44 of 77

Introduction to

AI Safety, Ethics, and Society

Normative Ethics

Part 3: Virtue Ethics and Social Contract Theory

44

Introduction to AI SES

45 of 77

Roadmap

  1. Why learn about ethics?

  1. Moral considerations: wellbeing, obligations, partiality, et cetera.

  1. Moral Theories:
    • Utilitarianism
    • Deontology
    • Virtue Ethics
    • Social Contract Theory

  1. Moral Uncertainty

45

Introduction to AI SES

46 of 77

Virtue Ethics

Virtue Ethics emphasizes the importance of having the right character traits.

46

Introduction to AI SES

Aristotle

Modern Virtue Ethics is inspired by the Ancient Greek philosopher, Aristotle. In his book, Nicomachean Ethics, Aristotle explored three key concepts:

  1. Virtue → good character traits.
  2. Practical Wisdom → skills required to live a good life.
  3. Flourishing → living a good life.

47 of 77

Virtue

To be virtuous is not just to behave in certain ways but also to feel certain ways.

Example. Bobby and Cory behave similarly: they are both trusted, keep their promises, and are equally honest. Bobby behaves virtuously because he feels its the right thing to do, but Cory behaves virtuously because she wishes to be seen as virtuous. To the virtue ethicist, only Bobby is virtuous.

Virtue ethics captures something important about morality that other theories neglect → mental states (emotions and motivations) are morally relevant.

Virtues are morally good character traits. Vices are morally bad ones.

47

Introduction to AI SES

48 of 77

Practical Wisdom

Practical wisdom is the ability to reason and act appropriately on the inclination to be virtuous.

If an individual lacks practical wisdom, the inclination to be virtuous may lead them to behave wrongly.

For instance, total honesty is not the ‘best policy’ – we should know when to be honest.

48

Introduction to AI SES

Even though people may be predisposed to certain virtues, practical wisdom – knowing how and when to act on such virtues – is gained through experience.

49 of 77

Flourishing

Flourishing, or eudaimonia, is living a good life.

To flourish, an individual needs to live virtuously, and virtues are those character traits which allow the individual to flourish.

Virtue ethicists still debate whether being virtuous is sufficient for living a good life.

49

Introduction to AI SES

Aristotle argued that to flourish, a virtuous individual must also have the resources to enact virtue.

50 of 77

Critiques of Virtue Ethics (1/2)

Virtue ethics fails for the precise reason many find it attractive. People who act virtuously despite not being virtuous are not considered praiseworthy.

Virtue ethics is not action-guiding. When faced with moral dilemmas, virtue ethics does not offer us a robust decision procedure we may follow in our moral judgments.

Virtue ethics is too focused on the individual. What is right or wrong depends on the character of the actor, which is odd when considering that ethics is concerned with how we treat others.

50

Introduction to AI SES

51 of 77

Social Contract Theory

51

Introduction to AI SES

Social contract theory focuses on hypothetical agreements between members of a society.

  • “Do not kill” → Social contract theorists might adopt this rule because it is in individual’s best interest not to kill each other.

Social contract theory views moral codes as the result of hypothetical agreements between members of society, established for mutual benefit. It claims that social contracts are the foundations of ethics.

Under this view, all moral codes are similarly justified.

52 of 77

The Veil of Ignorance

52

Introduction to AI SES

Developed by the contemporary moral theorist John Rawls, the veil of ignorance can be a decision-making tool for creating a social contract.

‘Behind’ the veil of ignorance, individuals lose all knowledge of their personal attributes: talents, religion, gender, sexuality, race, class, etc.

Once in this state, called the original position, individuals are asked to envision a social contract for society, blind to their own position within it.

Slavery. An individual behind the veil would not reasonably permit slavery given that they can’t know whether they are slave or slaver.

53 of 77

Protecting the Worst Off

53

Introduction to AI SES

Behind the veil of ignorance, group interests aren’t favored, since no one knows whether they belong to any given group.

  • If we don’t know which group we belong to, it’s possible that we end up in a position that is the “worst off”.

Thus, Rawls argues that everyone would ensure that the lowest level of wellbeing of anyone is sufficiently high.

Rawls’ Maximin Principle: society should prioritize maximizing the wellbeing of the person with the minimum level of wellbeing in society.

54 of 77

Protecting Liberty

54

Introduction to AI SES

Behind the veil of ignorance, any individual could potentially be excluded from having basic liberties.

  • If we don’t know whether we may/may not have certain liberties, we should distribute liberties equally to all individuals.

Individuals should be free to pursue their conception of the good life, and enjoy both civil and political liberties.

The Liberty Principle: the fair distribution of universal liberties, ensured by a contractarian agreement not to infringe upon the liberties of others.

55 of 77

Protecting Equality

55

Introduction to AI SES

Inequalities should arise only through fair access to opportunity for all.

Difference Principle: we accept inequality if there is equality of opportunity.

Inequalities should benefit the least privileged individuals - ‘worst off’.

  • Example: Economic Growth. Inequality can drive economic growth by rewarding productive people, which can also benefit the ‘worst off’.

Difference Principle (2): Inequalities arising from equality of opportunity must also help the least privileged, even if they help the privileged more.

56 of 77

Critique: Rawls’ Conclusions are Too Strong

56

Introduction to AI SES

The maximin principle is at odds with with common sense morality.

  • We should prioritize improving everyone’s wellbeing only if it also benefits the ‘worst off’ individual.

The maximin principle implies that the grouch should take priority.

  • The ‘worst off’ individual might just be extremely hard to help, such as someone comatose with an incurable illness. Why should we aim only to increase their wellbeing when more good could be done for others?

The principle seems to be untenable.

57 of 77

Rawls’ Conclusions Might Not Follow (1/2)

57

Introduction to AI SES

Behind the veil of ignorance, we might care about more than maximin.

Why would we always prioritize the “grouch”? Even for risk-averse individuals, ensuring a positive general distribution of wellbeing across society, rather than just maximizing average wellbeing, seems more reasonable.

Endorse the liberty and difference principles, we might not support maximin. If the ‘grouch’ has the same opportunities and liberties as anyone else, we might think we have no special reason to prioritize their wellbeing.

58 of 77

Rawls’ Conclusions Might Not Follow (2/2)

58

Introduction to AI SES

The veil of ignorance can be used to support utilitarianism. Decisions under uncertainty often involve maximizing expected/average results.

Economist John Harsanyi has argued that behind the veil of ignorance, rational agents would aim to maximize the total amount of wellbeing in a given society.

In experiments, people do not often choose maximin principles. Experiments simulating the veil of ignorance have found that participants tend to favor the ‘greater good’ or utilitarian outcomes.

John Harsanyi

59 of 77

Alternatives: Prioritarianism (1/3)

59

Introduction to AI SES

Rawls’ theory was a response to utilitarianism, but if the veil of ignorance favors utilitarian thinking, it could compromise the foundations of his theory.

Is there a middle ground?

Yes! → Prioritarianism is an ethical theory that attributes greater moral weight to improving the wellbeing of those who are worse off in society.

Crucially, this theory still takes into account the wellbeing of others.

60 of 77

Alternatives: Prioritarianism (2/3)

60

Introduction to AI SES

Imagine a situation where we have the option to distribute resources among three people, A, B, and C. We can choose how to change their wellbeing in three different ways.

A has 6 units of wellbeing.

B has 5 units of wellbeing.

C has 1 unit of wellbeing.

(+2) A has 8 units of wellbeing.

(+2) B has 7 units of wellbeing.

(+0) C has 1 unit of wellbeing.

(+1) A has 7 units of wellbeing.

(+1) B has 6 units of wellbeing.

(+1) C has 2 unit of wellbeing.

(–3) A has 3 units of wellbeing.

(–2) B has 3 units of wellbeing.

(+2) C has 3 unit of wellbeing.

1

2

3

61 of 77

Alternatives: Prioritarianism (3/3)

61

Introduction to AI SES

Which option should we choose?

Option 1 → maximizes wellbeing.

Option 2 → helps everyone a bit.

Option 3 → helps the disadvantaged person but hurts others.

Prioritarianism captures the intuition that (contra Rawls) option 3 seems worse than the others and (contra utilitarianism) option 1 is not clearly the best option.

Educational Policy. Prioritarians may focus their educational interventions on disadvantaged students whereas utilitarians may focus on policy interventions that benefit large numbers of students.

62 of 77

Alternatives: Moral Contractualism (1/2)

62

Introduction to AI SES

T.M. Scanlon’s moral contractualism suggests that people are inclined to seek out reasonable moral agreements.

Under this theory, Rawls’ veil of ignorance is unnecessary → morality is about what we owe each other as rational beings.

The principles of morality should be something that rational people can generally accept – they should be reasonable.

Example: It is reasonable to reject moral codes that support slavery.

T.M. Scanlon

63 of 77

Recap

63

Introduction to AI SES

We have thus far outlined the four most common ethical theories:

  1. Utilitarianism → ethical decisions aim to maximize wellbeing.
  2. Deontology → ethical decisions are in made in accordance with the right system of rules.
  3. Virtue Ethics → ethical decisions come from a virtuous character.
  4. Social Contract Theory → ethical decisions are made in accordance with with mutually-agreed upon moral principles.

Each theory focuses on different kinds of moral considerations that we often consider when thinking about “common-sense morality”. This illustrates how different moral theories can complement each other.

64 of 77

Introduction to

AI Safety, Ethics, and Society

Normative Ethics

Part 4: Moral Uncertainty

64

Introduction to AI SES

65 of 77

Roadmap

  1. Why learn about ethics?

  1. Moral considerations: wellbeing, obligations, partiality, et cetera.

  1. Moral Theories:
    • Utilitarianism
    • Deontology
    • Virtue Ethics
    • Social Contract Theory

  1. Moral Uncertainty

65

Introduction to AI SES

66 of 77

Decisions Under Moral Uncertainty

How do we act morally when we are unsure which moral view is correct? We might have different credences in different moral theories.

Credence: varying degrees of belief in different moral theories. It’s often expressed as a probability value an individual assigns to any given theory.

  • Ann may have a high degree of credence in utilitarianism – 70% – and a low degree in deontology – 30%.

One approach is to follow a reasonable pluralism, acknowledging the potential co-existence of multiple reasonable moral theories.

66

Introduction to AI SES

67 of 77

Dealing with Moral Uncertainty

67

Introduction to AI SES

In high-stakes scenarios, a reasonable pluralism is insufficient. We may need to consider concrete ethics over conflicting wisdom that appears reasonable.

In AI, the stakes are high. The implications of AI systems are widespread, affecting aspects of our lives from healthcare to security.

Accounting for moral uncertainty can help us avoid bad outcomes. It is crucial that AIs behave in ways that correspond with human values and goals – recognizing moral uncertainty is vital to this process.

68 of 77

How to Approach Moral Uncertainty?

68

Introduction to AI SES

There are several potential solutions to the problem of moral uncertainty. Let’s consider three:

  1. My Favorite Theory (MFT): adopting a favored moral theory that aligns with my personal beliefs.
  2. Maximize Expected Choice Worthiness (MEC): calculating the expected choice worthiness of each option to obtain an average moral value.
  3. Moral Parliament: treating a moral decision as a negotiation in a parliament, considering multiple moral views to find a mutually acceptable solution.

69 of 77

How to Approach Moral Uncertainty?

69

Introduction to AI SES

Example: Should we save a life?

Imagine a notorious murderer asks Alex where his friend, Jordan, is. Alex knows that revealing Jordan's location will likely lead to his friend's death. However, while lying would save his life, it is morally questionable.

Alex is unsure about which action to take and considers the recommendations of the three moral theories he has some credence in: Utilitarianism (60%), Deontology (30%), and Contractarianism (10%).

Alex must decide which action to take, given the varied perspectives of these moral theories.

70 of 77

My Favorite Theory

70

Introduction to AI SES

Under MFT, Alex would pick whatever Utilitarianism recommends.

Upside → this approach is simple and easy to implement, not requiring an understanding of multiple different theories.

Downside → it only works when the level of moral uncertainty is low.

  • Following MFT can lead to single-mindedness or overconfidence

  • MFT can discard relevant information, such as when the credences in two theories are close, but their judgments of an action vastly differ.

71 of 77

MEC (1/3)

71

Introduction to AI SES

Under MEC, Alex must figure out choice worthiness for each theory:

  • Utilitarianism values lying to save a life highly, so Alex assigns it a value of +500.

  • Deontology strongly disapproves of lying, even to save a life, so Alex assigns it a value of -1000.

  • Contractarianism moderately approves of lying to save a life, so Alex assigns it a value of +100.

72 of 77

MEC (2/3)

72

Introduction to AI SES

Next, Alex needs to multiply choice-worthiness by credence.

We know that Alex has 60% credence in Utilitarianism, 30% credence in Deontology, and 10% credence in Contractarianism.

Therefore, Alex does the following calculation:

.6(500) + .3(-1000) + .1(100) =

300 - 300 +10 =

10

Under MEC, Alex would choose to lie since it is the best average moral outcome.

73 of 77

MEC (3/3)

73

Introduction to AI SES

However, MEC faces problems with comparisons between theories.

The assignment of choice-worthiness can be arbitrary, and with ‘ordinal’ theories that rank actions by their moral value, determining choice-worthiness can be difficult.

MEC also cannot account for absolutist theories – we assigned a value of -1000 to lying under deontology, but it is unclear whether this actually reflects the absolute value of lying. A more accurate choice might have been negative infinity, which would swamp all else.

74 of 77

Moral Parliament (1/2)

74

Introduction to AI SES

Following a Moral Parliament approach, Alex could more carefully consider the proposal to lie by assigning imagined delegates to each view, with 60 for Utilitarianism, 30 for Deontology, and 10 for Contractarianism.

Depending on the votes, Moral Parliament can lead to different outcomes. �

  • Majority rule: Utilitarians will likely have the final say.
  • Philosophers recommend using proportional chance voting instead, where an action is taken with a probability equal to its vote-share.
  • If no one changed their mind, the outcome of the parliament would recommend with 70% probability that Alex lie and with 30% probability that he tells the truth.

75 of 77

Moral Parliament (2/2)

75

Introduction to AI SES

Even though Moral Parliament encourages negotiation, it is difficult to see what it might recommend. One question is how the theoretical delegates actually change their mind or compromise.

The outcome of the Moral Parliament is not determined externally – it is a matter of our imagination, and subject to our biases.

Moral Parliament may not be very helpful for individuals thinking about what is right to do in practice—although it might be good for AIs!

76 of 77

Recap

76

Introduction to AI SES

Resolving moral uncertainty is difficult but important.

In high stakes scenarios like AI design and development, addressing moral uncertainty is crucial.

Three solutions to moral uncertainty are:

  1. My Favorite Theory
  2. Maximize Expected Choice Worthiness
  3. Moral Parliament

There might not be a one-size-fits-all solution to all ethical dilemmas.

77 of 77

Chapter Recap

77

Introduction to AI SES

Ethics is the study of moral principles that guide decisions and actions.

AI researchers can use ethics to guide development and governance.

Moral considerations include goodness, constraints, special obligations, and more.

Utilitarianism focuses on consequences and wellbeing; deontology on constraints; virtue ethics on character, and social contract theory on agreements.

We can deal with uncertainty in many ways, such as moral parliaments.