I'm a Philosophy Fellow at the Center for AI Safety. I've been thinking about coherence, corrigibility, and myopia, among other things.
Before that, I was a PhD student in Philosophy and Parfit Scholar at the Global Priorities Institute. My thesis is in population ethics. I've also done some thinking about decision theory, moral uncertainty, and cost-benefit analysis.
You can email me at [email protected]
Thanks, Danny! This is all super helpful. I'm planning to work through this comment and your BCA update post next week.
I think this paper is missing an important distinction between evolutionarily altruistic behaviour and functionally altruistic behaviour.
These two forms of behaviour can come apart.
A parent's care for their child is often functionally altruistic but evolutionarily selfish: it is motivated by an intrinsic concern for the child's welfare, but it doesn't confer a fitness cost on the parent.
Other kinds of behaviour are evolutionarily altruistic but functionally selfish. For example, I might spend long hours working as a babysitter for someone unrelated to me. If I'm purely motivated by money, my behaviour is functionally selfish. And if my behaviour helps ensure that this other person's baby reaches maturity (while also making it less likely that I myself have kids), my behaviour is also evolutionarily altruistic.
The paper seems to make the following sort of argument:
I think we have reasons to question premises 1 and 2.
Taking premise 2 first, recall that evolutionarily selfish behaviour can be functionally altruistic. A parent’s care for their child is one example.
Now here’s something that seems plausible to me:
If that’s the case, then functionally altruistic behaviour is evolutionarily selfish for AIs: this kind of behaviour confers fitness benefits. And functionally selfish behaviour will confer fitness costs, since we humans are more likely to shut off AIs that don’t seem to have any intrinsic concern for human welfare.
Of course, functionally selfish AIs could recognise these facts and so pretend to be functionally altruistic. But:
Here’s another possible objection: functionally selfish AIs can act as a kind of Humean ‘sensible knave’: acting fairly and honestly when doing so is in the AI’s interests but taking advantage of any cases where acting unfairly or dishonestly would better serve the AI’s interests. Functionally altruistic AIs, on the other hand, must always act fairly and honestly. So functionally selfish AIs have more options, and they can use those options to outcompete functionally altruistic AIs.
I think there’s something to this point. But:
Here’s another possible objection: AIs that devote all their resources to just copying themselves will outcompete functionally altruistic AIs that care intrinsically about human welfare, since the latter kind of AI will also want to devote some resources to promoting human welfare. But, similarly to the objection above:
Okay, now moving on to premise 1. I think you might be underrating group selection. Although (by definition) evolutionarily selfish AIs outcompete evolutionarily altruistic AIs with whom they interact, groups of evolutionarily altruistic AIs can outcompete groups of evolutionarily selfish AIs. (This is a good book on evolution and altruism, and there’s a nice summary of the book here.)
What’s key for group selection is that evolutionary altruists are able to (at least semi-reliably) identify other evolutionary altruists and so exclude evolutionary egoists from their interactions. And I think, in this respect, group selection might be more of a force in AI evolution than in biological evolution. That’s because (it seems plausible to me) that AIs will be able to examine each other’s source code and so determine with high accuracy whether other AIs are evolutionary altruists or evolutionary egoists. That would help evolutionarily altruistic AIs identify each other and form groups that exclude evolutionary egoists. These groups would likely outcompete groups of evolutionary egoists.
Here’s another point in favour of group selection predominating amongst advanced AIs. As you note in the paper, groups consisting wholly of altruists are not evolutionarily stable, because any egoist who infiltrates the group can take advantage of the altruists and thereby achieve high fitness. In the biological case, there are two ways an egoist might find themselves in a group of altruists: (1) they can fake altruism in order to get accepted into the group, or (2) they can be born into a group of altruists as the child of two altruists, and (by a random genetic mutation) can be born as an egoist.
We already saw above that (1) seems less likely in the case of AIs who can examine each other’s source code. I think (2) is unlikely as well. For reasons of goal-content integrity, AIs will have reason to make sure that any subagents they create share their goals. And so it seems unlikely that evolutionarily altruistic AIs will create evolutionarily egoistic AIs as subagents.
I wouldn't call a small policy like that 'democratically unacceptable' either. I guess the key thing is whether a policy goes significantly beyond citizens' willingness to pay not only by a large factor but also by a large absolute value. It seems likely to be the latter kinds of policies that couldn't be adopted and maintained by a democratic government, in which case it's those policies that qualify as democratically unacceptable on our definition.
Yes, I think so!
And thanks again for making this point (and to weeatquince as well). I've written a new paragraph emphasising a more reasonable, less conservative estimate of benefit-cost ratios. I expect it'll probably go in the final draft, and I'll edit the post here to include it as well (just waiting on Carl's approval).
I think this is right (and I must admit that I don't know that much about the mechanics and success-rates of international agreements) but one cause for optimism here is Cass Sunstein's view about why the Montreal Protocol was such a success (see Chapter 2): cost-benefit analysis suggested that it would be in the US's interest to implement unilaterally and that the benefit-cost ratio would be even more favourable if other countries signed on as well. In that respect, the Montreal Protocol seems akin to prospective international agreements to share the cost of GCR-reducing interventions.
Thanks for this! All extremely helpful info.
This is good to know. Our BCR of 1.6 is based on very conservative assumptions. We were basically seeing how conservative we could go while still getting a BCR of over 1. I think Carl and I agree that, on more reasonable estimates, the BCR of the suite is over 5 and maybe even over 10 (certainly I think that's the case for some of the interventions within the suite). If, as you say, many people in government are looking for interventions with BCRs significantly higher than 1, then I think we should place more emphasis on our less conservative estimates going forward.
Thanks very much for this! I might try to get some of these references into the final paper.
This is really good to know as well.
I think this is right. Our claim is that a strong longtermist policy as a whole would place extreme burdens on the present generation. We expect that a strong longtermist policy would call for particularly extensive refuges (and lots of them) as well as the other things that we mention in that paragraph.
We use that threshold because we think that focusing on that threshold by itself makes the benefit-cost ratio come out greater than 1. I’m not so sure that’s the case for the more common thresholds of killing at least 1 billion people or at least 10% of the population in order to qualify as a global catastrophe.
We're not opposed to including effects on future generations in cost-benefit calculations. We do the calculation that excludes benefits to future generations to show that, even if one totally ignores benefits to future generations, our suite of interventions still looks like it's worth funding.
Oh interesting! Thanks.
And thanks very much for this! I think we will still be able to mention this in the published version.
Yep, agreed!
My impression is that this is being explored. See, e.g., here.
The argument we mean to refer to here is the one that we call the ‘best-known argument’ elsewhere: the one that says that the non-existence of future generations would be an overwhelming moral loss because the expected future population is enormous, the lives of future people are good in expectation, and it is better if the future contains more good lives. We think that this argument is liable to overshoot.
I agree that there are other compelling longtermist arguments that don’t overshoot. But my concern is that governments can’t use these arguments to guide their catastrophe policy. That’s because these arguments don’t give governments much guidance in deciding where to set the bar for funding catastrophe-preventing interventions. They don’t answer the question, ‘By how much does an intervention need to reduce risks per $1 billion of cost in order to be worth funding?’.
This seems like a good target to me, although note that $400b is our estimate for how much it would cost to fund our suite of interventions for a decade, rather than for a year.
Yes, this is an important point. If we were to do a more detailed cost-benefit analysis of catastrophe-preventing interventions, we’d want to address it more comprehensively (especially since we also mention how different interventions can undermine each other elsewhere in the paper).
On the point about the average only just meeting the bar, though, I think it’s worth noting that our mainline calculation uses very conservative assumptions. In particular, we assume:
And we count only these interventions benefits in terms of GCR-reduction. We don’t count any of the benefits arising from these interventions reducing the risk of smaller catastrophes.
I think that, once you replace these conservative assumptions with more reasonable ones, it’s plausible that each intervention we propose would pass a CBA test.
I think that these points also help our conclusions apply to other countries, even though all other countries are either smaller than the US or employ a lower VSL. And even though reasonable assumptions will still imply that some GCR-reducing interventions are too expensive for some countries, it could be worthwhile for these countries to participate in a coalition that agrees to share the costs of the interventions.
I agree that speculative estimates are a major problem. Making these estimates less speculative – insofar as that can be done – seems to me like a high priority. In the meantime, I wonder if it would help to emphasise that a speculative-but-unbiased estimate of the risk is just as likely to be too low as to be too high.
I agree with this. What we’re arguing for is a criterion: governments should fund all those catastrophe-preventing interventions that clear the bar set by cost-benefit analysis and altruistic willingness to pay. One justification for funding these interventions is the justification provided by CBA itself, but it need not be the only one. If longtermist justifications help us get to the place where all the catastrophe-preventing interventions that clear the CBA-plus-AWTP bar are funded, then there’s a case for employing those justifications too.
We think that longtermists in the political sphere should (as far as they can) commit themselves to only pushing for policies that can be justified on CBA-plus-AWTP grounds (which need not entirely ignore effects on future generations). We think that, in the absence of such a commitment, the present generation may worry that longtermists would go too far. From the paper: