Evan Hubinger (he/him/his) ([email protected])
I am a safety researcher at Anthropic. My posts and comments are my own and do not represent Anthropic's positions, policies, strategies, or opinions.
Previously: MIRI, OpenAI
See: “Why I'm joining Anthropic”
Selected work:
I think this is extremely not true, and am pretty disappointed with this sort of "debate me" communications policy. In my opinion, I think public debates very rarely converge towards truth. Lots of things sound good in a debate but break down under careful analysis, and the pressure of saying things that look good to a public audience creates a lot of pressure opposed to actual truth-seeking.
I understand and agree with the importance of good communications here, but imo this is really not the way. Some alternative possibilities:
I'm sure there's a bunch more here; these are just some ideas off the top of my head. In general, I think there's a lot of ways to do public communications on complex, controversial topics that don't involve public debates, and I'd strongly encourage going in one of those alternative directions instead.
Cross-posted from LessWrong.
In my opinion, I think the best solution here is incentivizing people to voluntarily have more children—e.g. child tax credits, maternity/paternity leave, etc. If you don't think fetuses are moral patients, then the pro-natalist, longtermist, total utilitarian view doesn't distinguish between having an abortion and just choosing not to have a child, so I don't really see the reason to focus on abortion specifically in that case.
I am absolutely intending to communicate that I think it would be good for people to say that they think fraud is bad. But that doesn't mean that I think we should condemn people who disagree regarding whether saying that is good or not. Rather, I think discussion about whether it's a good idea for people to condemn fraud seems great to me, and my post was an attempt to provide my (short, abbreviated) take on that question.
I don't think this and didn't say it. If you have any quotes from the post that you think say this, I'd be happy to edit it to be more clear, but from my perspective it feels like you're inventing a straw man to be mad at rather than actually engaging with what I said.
I think that, for the most part, you should be drawing your ethical boundaries in a way that is logically prior to learning about these sorts of facts. Otherwise it's very hard to cooperate with you, for example.
It isn't intended as a rhetorical question. I am being quite sincere there, though rereading it, I see how you could be confused. I just edited that section to the following:
I actually state in the post that I agree with this. From my post:
Perhaps that is not as clear as you would like, but like I said it was a short post. And that sentence is pretty clearly saying that I think it's worthwhile for us to try to carefully confront the moral question of what is okay and what is not—which the post then attempts to start the discussion on by providing some of what I think.
I guess I don't really think this is a problem. We're perfectly comfortable with statements like “murder is wrong” while also understanding that “but killing Hitler would be okay.” I don't mean to say that talking about the edge cases isn't ever helpful—in fact, I think it can be quite useful to try to be clear about what's happening on the edges in certain cases, since it can sometimes be quite relevant. But I don't see that as a reason to object to someone saying “murder is wrong.”
To be clear, if your criticism is “the post doesn't say much beyond the obvious,” I think that's basically correct—it was a short post and wasn't intended to accomplish much more than basic common knowledge building around this sort of fraud being bad even when done with ostensibly altruistic motivations. And I agree that further posts discussing more clearly how to think about various edge cases would be a valuable contribution to the ongoing discussion (though I don't personally plan to write such a post because I think I have more valuable things to do with my time).
However, if your criticism is “your post says edge case B is bad but edge case B is actually good,” I think that's a pretty silly criticism that seems like it just doesn't really understand or engage with the inherent fuzziness of conceptual categories.
It's clearly not fully general because it only applies to excluding edge cases that don't satisfy the reasons I explicitly state in the post.
Sure, but that's not what happened. There are some pretty big disanalogies between the scenarios you're describing and what actually happened:
Adding on to my other reply: from my perspective, I think that if I say “category A is bad because X, Y, Z” and you're like “but edge case B!” and edge case B doesn't satisfy X, Y, or Z, then clearly I'm not including it in category A.
I think you're wrong about how most people would interpret the post. I predict that if readers were polled on whether or not the post agreed with “lying to Nazis is wrong” the results would be heavily in favor of “no, the post does not agree with that.” If you actually had a poll that showed the opposite I would definitely update.
I think my post is quite clear about what sort of fraud I am talking about. If you look at the reasons that I give in my post for why fraud is wrong, they clearly don't apply to any of examples of justifiable lying that you've provided here (lying to Nazis, doing the least fraudulent thing in a catch-22, lying by accident, etc.).
In particular, if we take the lying to Nazis example and see what the reasons I provide say:
This clearly doesn't apply to lying to Nazis, since it's not a situation where money and power are being seized for oneself.
I think the fact that you would lie to a Nazi makes you more trustworthy for coordination and cooperation, not less.
And in the case of lying to Nazis, the consequences are clearly positive.