AMA: Anthony DiGiovanni, author of the ‘Challenge of Unawareness’ sequence

By Toby Tremlett🔹, Anthony DiGiovanni 🔸 @ 2026-07-03T09:14 (+39)

We announced the Cluelessness Critiques Competition two weeks ago. 

A lot of you, not only prospective entrants, will have been reading Anthony's sequence where he lays out his unawareness argument, or following the comments on his summary post. I thought that this might be a great time to have Anthony put some time aside to answer your questions.

Although we are calling this an AMA[1], the focus will be on helping people understand the sequence so that they can write the best entries to the competition that they can. Anthony will be choosing the questions he responds to with this in mind.

Anthony will be answering your questions on Thursday the 9th. He cannot guarantee that he will answer every question, so make sure to upvote the questions you’d like to see answered. 

  1. ^

    Which stands for 'Ask Me Anything'


Ben_West🔸 @ 2026-07-03T22:09 (+8)

You respond to Richard Ngo here:

> do you think that, if we had a theory of sociopolitics that was about as good as 20th-century economics, then we wouldn't be clueless about how to do sociopolitical interventions (like founding AI safety movements) effectively?

No, because I think “founding AI safety movements that succeed at making the far future go better” is a pretty out-of-distribution kind of sociopolitical intervention.

Suppose instead we had a comparably good theory of the right reference class, e.g. "movements trying to shape transformative technologies." Would we still be clueless about AI safety movement-building? 

More generally: you list various considerations across your posts and I have a hard time understanding which is load-bearing for your answer here. Some possibilities: 

  1. We're clueless because we haven't yet developed the relevant theory (Richard's reading IIUC, on which cluelessness is contingent and reducible)
  2. No such theory could be validated even in principle, because we never observe the target variable (far-future value) and calibration on near-term proxies doesn't transfer
  3. Even a validated theory wouldn't help, because impact is dominated by considerations inaccessible to any theory (e.g. unconceived hypothesis classes)
Anthony DiGiovanni 🔸 @ 2026-07-09T16:08 (+4)

I think it's mostly (1), but I'm open to something like (2) or (3) as well.

(Following (1):) There is in principle some (a) amount of information that non-ideal agents could attain about the cosmos with non-Pascalian probability,[1] + (b) a priori modeling and induction we could apply to that information, such that we wouldn't be clueless. So I don't think we need to observe the target variable, or empirically "validate" the theory, to be non-clueless.

But the bar to achieve such an (a)+(b) seems very high, because:

  • If we do try to empirically validate the theory by appealing to calibration on near-term proxies:
    • I indeed don't see why we should expect such calibration to transfer, up to the degree of precision we need to escape cluelessness (sec. 2.3.1.1). This bites even if, say, we use AI to get much more calibrated on ~years-long time horizons.
  • If we don't, and instead try to argue conceptually that the theory captures enough of the relevant considerations in fine-grained enough detail:
    • The web of factors this theory would have to capture seems ludicrously complex (the rest of sec. 2.3). Of course, good theories can compress complexity, but getting that amount of compression while keeping things computationally tractable[2] sounds rough.

So my suspicion is that yeah, we'd still be clueless given the kind of theory you mention. But I find it hard to say, because I can't imagine exactly what "comparably good" looks like, concretely. I appreciate that that's hard to spell out on your end.

Maybe sufficiently advanced AI could get around this. Maybe not, e.g. if "the universal prior" is irreducibly imprecise, or if (following (3)) information about simulators or causally disconnected worlds is fundamentally inaccessible.

(I'm happy to unpack any of this more if useful, not sure if I answered your question properly!)

  1. ^

    As in, if I were to represent this probability numerically, the interval wouldn't all be less than the Pascalian threshold.

  2. ^

    Like, something analogous to the Schrödinger equation doesn't count. :)

Ben_West🔸 @ 2026-07-03T20:24 (+8)

What's the clearest example of a complex cluelessness sign flip you're aware of?

(By "clear" I mean "had a very narrow confidence interval before encountering some consideration and a narrow interval after encountering that consideration but the CIs now center points with opposite signs".[1])

The clearest examples I know of (e.g. rescuing Hitler as a child) seem to me like examples of simple cluelessness. You list some examples here, but they don't seem that clear to me, e.g. I disagree that "Early awareness-raising about AGI x-risk presumably seemed robustly good" and would guess most people involved in that had CIs which comfortably straddled zero. 

  1. ^

    Or alternatively: there are two representors with narrow but non-overlapping CIs.

Anthony DiGiovanni 🔸 @ 2026-07-09T08:21 (+3)

Hi Ben, I like the spirit of this question, though I'm not sure it's the most relevant formulation. Thoughts on that, before I answer your literal question:

  • To get clear on terms:
    • I'm assuming by "CI" here, you mean something like your precise 90% (or whatever) confidence interval for your idealized self's EV of the intervention (as per Premise 1).
    • By "robust", I don't mean a narrow / strictly positive CI. I mean that the verdict "this intervention is positive 'in expectation'" isn't sensitive to arbitrary choices about how to factor in the considerations we're unaware of.
      • I think it's fair to say most people involved in early AI risk advocacy considered their work robustly positive in that sense.
  • This definition of "robust" is what matters for Premise 3. Because P3 says, our verdicts about interventions having positive "EV" are sensitive to such arbitrary choices — even if we admit we're very uncertain (i.e. we have wide CIs).
  • So, if we're asking whether a given "sign flip" counts as evidence for P3, I don't see why the bar should be "narrow CIs with opposite-sign center points before and after the consideration".
    • If we've discovered a consideration that flipped us from "wide CI centered at a positive EV" to "wide CI centered at a negative EV", isn't that some evidence that our initial "positive in expectation" verdict wasn't robust, in the sense above? (And hence inductive evidence that our current "positive in expectation" verdicts aren't robust (i.e., evidence for P3), as argued here.)

Anyway, I'd agree that the clearest evidence for P3 would come from sign flips that meet your bar. Maybe the small animal replacement problem? I'd guess lots of people who care about animal welfare thought that getting people to eat less beef was clearly good before being aware of SARP, and think it's clearly net-bad after being aware of SARP. (It's harder to come by examples of sign flips by your def for longtermist causes, because non-clueless longtermists typically agree that we should be very uncertain about the far future. But per the above, this is to be expected if we're clueless.)

Ben_West🔸 @ 2026-07-09T21:21 (+11)

Thanks! I can't tell if this is cruxy, but for what it's worth your "pessimal induction" vignettes don't resonate with me in a way which makes me less motivated by the unawareness concerns.

For example, Bostrom coined the phrase "attention hazard" in 2011. I remember someone telling me that MIRI was net-negative for this reason at EAG 2015, and I would be surprised if e.g. Habryka hadn't considered this risk before starting Lightcone. So I disagree with citing him/this as a good example of unawareness; it's more that they mis-estimated a known risk factor. 

Similarly, I remember talking about SARP at one of my first EAGs. I think I came across it in Brian Tomasik's 2007 post, maybe even before I had encountered EA. Perhaps I've mis-estimated those concerns, but it doesn't seem like unawareness.

My overall experience is kind of the opposite of yours: when I got involved in EA people talked a lot about "Cause X" and "Crucial Considerations" and now they've mostly just... stopped? Like people tried to find other considerations, and there's some new stuff around s-risks and weird decision theories etc., but if you look at what people talk about at EAGs today it feels mostly like more precise versions of what was discussed in 2016, rather than a large and unpredictable jump from the older understanding. Or, more technically: it feels like we've had updates in evidence-space, but not as many updates in hypothesis-space, and I understand the latter to be motivating imprecision.

Obviously, this could be because EAs suck at cause prio research, or we just haven't been hit yet with the big update, etc., but the "pessimal induction" seems less pessimal to me.

Anthony DiGiovanni 🔸 @ 2026-07-11T16:57 (+2)

Interesting, that's helpful to know.

Not a comprehensive reply, but: I think many of the examples you're talking about are arguably cases of coarse awareness. People were coarsely aware of the potential backfire risks earlier on, but (arguably) the reason they didn't give these risks enough weight was that they didn't have a more fine-grained awareness of the specific causal pathways. I think such cases count as evidence for the pessimistic induction.

mal_graham🔸 @ 2026-07-11T14:30 (+4)

Sorry I’m probably missing something, but I’m not understanding why real world examples from EA would be particularly relevant given how young a movement it is. I think someone could grant that we have the ability to be justified in assigning probabilities to things that are likely to happen soon, and agree that the risk of things we’re totally unaware of happening in the next ~ 10-50 years might be (at least in some circumstances) sufficiently small to not have unawareness problems.

But once you start trying to be an impartial altruist about far future beings, that seems to me where you really can’t get away from unawareness problems. And so I guess if you wanted to convince me I was wrong about that, we should be looking at things that people thought 1000 years ago, and how things they caused today were bad even though they were trying to do good for reasons they weren’t only poorly calibrated on but in fact totally unaware of - and it just seems likely to me there would be tons of examples of that?

Maybe the development of gunpowder stands out here as something being pursued in the hopes of achieving eternal life (ostensibly an altruistic motivation) and presumably the possibility of guns was not on people’s radar. I guess it would eventually have been figured out anyway, but how much harm did having gunpowder X years earlier cause?

Maybe an objection here is that an “ideal” agent would have of course considered the possibility of any chemical work being misused, but IDK - they weren’t even trying to make something explosive. I don’t see how even a perfectly rational being could have predicted all the harms gunpowder would cause given that they were aiming to do alchemy. What probability could they have possibly been justified, given their epistemic position, in assigning to ”super bad outcomes from pursuing eternal life chemistry” given that they probably could not have imagined the scale of modern warfare?

I do get a little mixed up on this between “people are not ideal and so regularly make large mistakes that look like cluelessness” vs “even an ideal agent could not be justified in their probability assignments given what is theoretically knowable” so maybe I’m misunderstanding something.

Ben_West🔸 @ 2026-07-11T17:08 (+2)

Anthony cites Greaves and MacAskill giving an example similar to your gunpowder one:

Consider, for example, would-be longtermists in the Middle Ages. It is plausible that the considerations most relevant to their decision – such as the benefits of science, and therefore the enormous value of efforts to help make the scientific and industrial revolutions happen sooner – would not have been on their radar. Rather, they might instead have backed attempts to spread Christianity, perhaps by violence: a putative route to value that, by our more enlightened lights today, looks wildly off the mark. The suggestion, then, is that our current predicament is relevantly similar to that of our medieval would-be longtermists.

I personally think these examples are less compelling than they first appear (e.g. the persistence literature generally finds weaker effects than what you might imagine), but I agree that a failure of EAs to find examples of sign flips doesn't mean that future ones won't exist. 

mal_graham🔸 @ 2026-07-11T14:15 (+3)

Does P1 rely in any way on the idealized agent actually using EV specifically, as opposed to other theoretically possible aggregations of all or most of the possible consequences of A and B? It seems like no to me but I was curious since it is explicitly mentioned.

Anthony DiGiovanni 🔸 @ 2026-07-11T17:02 (+2)

Good question! I think "other theoretically possible aggregations of all or most of the possible consequences of A and B" would also suffice, yeah. (Of course, if we ourselves can't specify what this alternative is, we have our work cut out for us if we're gonna argue that we should expect our idealized self to prefer A over B on this basis.)

mal_graham🔸 @ 2026-07-11T14:11 (+3)

Can you explain a little more what you mean by “coarse-grained” in P3? Does that just mean “very unlikely to include all the possible outcomes, so there’s a ton of unknown unknowns we don’t know how to assign probability to” or something else?

Anthony DiGiovanni 🔸 @ 2026-07-11T17:24 (+2)

I mean that we have what I call "coarse awareness" here: we conceive of crude groups of possible worlds, rather than possible worlds specified in fine-grained enough detail to assign them precise values (wrt impartial altruist axiologies). See also here for some examples. Happy to unpack more if those sections don't answer things!

dan.pandori 🔸 @ 2026-07-09T00:13 (+3)

Do you see cluelessness to be decreasing, steady, or increasing?

As in, are we getting better are predicting longterm outcomes or do you think we are no better than centuries ago?

One vote for decreasing cluelessness: improved epistemic practices such as RCTs, peer review, more advanced statistics, and prediction markets. There used to be many actions which we were deeply unsure about the near-term consequences of (ex. is bleeding this patient a good idea), and now our best predictions seem justifiably more confident. The evidence for improved long-term predictions seems less obvious, but still directionally improving.

Anthony DiGiovanni 🔸 @ 2026-07-09T16:16 (+2)

I'd guess we're getting slightly better, yep. I might put less weight on the evidence you mention, than on: "We're living in a period of really unprecedented AI progress, seems like that puts in a better position to reason about the mechanisms governing the far future than ever before."