What I want you to do when I tell you to “think about your theory of change more carefully”

By Roman Ross @ 2026-09-12T02:48 (+6)

Summary: I worry that a lot of people enter into projects because they “seem good,” when actually, they miss a lot of important steps necessary for making a project that’s going to be especially impactful. This post argues that writing out a “theory of change” is important and gives examples of how to do this well.  

Tell the full story

One thing that makes AI safety stand out from other EA cause areas is that it’s hard to directly measure the impact of various interventions. When comparing global health outcomes, you can measure the cost-effectiveness and impact of your treatment with RCTs. When estimating how many chicken-days you can affect by running a cage-free campaign, you can analyze the impact of similar past actions. However, in AI safety, we don’t have as many helpful feedback loops to see how good our work is. And while you might come up with some proxy for this, you can risk Goodharting yourself if you don’t understand your analytics properly.

To get around this, we try to make sure that our interventions have coherent and strong “theories of change”: detailed explanations of what needs to happen for a project to have a good impact on the world. Here’s what I mean by that:

Some examples of theories of change that I think are underdeveloped and weak:

Some examples of theories of change that I think are much better:

The theories of change in the second group are better not because they are longer, but because they “backchain” from the ultimate goal. This means that they have a clear vision of what “winning” looks like in AI safety, and what further steps are needed to take us there. How does your project fill in one of these critical steps? This also means you should have thought carefully about what the future could look like and have a developed “theory of victory.”

Find a metric

One trick that’s helpful for determining your theory of change is finding a metric you could use to calculate your impact. This forces you to tell a more coherent story about where the impact of your project comes from, and it helps you escape from traps where you repeatedly tell yourself, “My theory of change is that it’s good for knowledge to exist!” For example:

For this trick to help, it doesn’t really matter if your metric is somewhat impossible to measure, or if your units are especially realistic. You’re not aiming to do a BOTEC, you’re aiming to think clearly about your goals.

I don’t think this trick is universally applicable. For example, I’m not sure how I would create a metric to measure the impact of policy work. I might try “number of people now willing to consider a pause on AI development a reasonable proposal” or “increase in lab willingness-to-pay for safety,” but these are confusing and illegible. Despite this, I still think that there are lots of other cases where developing metrics can help clarify your thinking.

Lots of things “seem good” but are missing important steps to becoming impactful

A common mistake I see people make is stopping a project at a point where an additional ~30% of its value is easily capturable with a few more additions/changes. Here are some examples:

Understanding your project’s theory of change can be helpful for remembering that these steps are very important. 

Lots of things “seem good” but are actually worse than nearby options

Sometimes, thinking about your project in depth can make you realize there’s a better way to do the same thing. Examples:

Things to avoid when thinking about theories of change

  1. Don’t worry too much about this if you expect the thing you’re doing to take less than two hours.
  2. Don’t overthink things too much. You can imagine reasons why anything might fail if you try hard enough, and you shouldn’t drive yourself into paralysis with skepticism. 
  3. If you’re in a fieldbuilding role, thinking too much about pushing people through “talent pipelines” can be harmful toward your goals. I’m uncertain about the best way to navigate this issue, but I don’t think you should treat/think about other people only as a means to an end.

Final Note:

If you think this post is missing anything or gets anything wrong, please say so in a comment. I intend to link this post to a bunch of people I talk to as a way of saving time, so your input will probably be heard by many people who would especially benefit from hearing it. 

  1. ^

    I’m not a technical researcher, so forgive me if this guess is horribly off-base.

  2. ^

    This isn’t to say that EQ Bench 4 is completely meaningless or unhelpful, only that I can’t easily figure out what it means or why it helps me.

  3. ^

    A counter-intuition is that recommendations from friends are the strongest form of advertising, and you only get that if your product is really, really good. Still, if advertising is as cheap as making a LinkedIn post or dm-ing it to some people who might benefit from having it, advertising is probably worth it.