Community Polls on Alignment Controversies II

By Jasmine Brazilek, Jonah Woodward, Miles Tidmarsh @ 2026-07-30T20:25 (+30)

Please spend <5 minutes filling in the below polls on AI alignment!

Thank you to everyone who filled out last month's polls. It was great to see 60+ comments engaging with these issues.

This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We'll compare panel and community responses in an upcoming report, which we'll publish here and on LessWrong. To get notified when it's released, you can subscribe to our new Substack.

Many people we've talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML's research agenda.

A few final things about the polls themselves:

Thanks to BlueDot Impact for funding this work.

 
 
 
 
 
 
 
 
 
 
 
 
 
 

 

 

 

¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation. 

² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted. 

³ This is about where the next dollar is best spent, not about which area you think is more important overall.

⁴ "Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.

⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.

⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.

⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first). 

8 AIs that are trying to hide features of themselves from humans and operators

*By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.


toasty_sunbeam @ 2026-08-03T20:55 (+3)

If animals continue to exist in a post-AGI world, animal suffering will not persist

Suffering is part of the ecosystem. The only way to end animal suffering altogether would be to wipe out all animal life. If the question was "most animal suffering", I would have a different answer.

Vakus Drake @ 2026-08-03T21:54 (+2)

You aren't thinking enough about the possibilities enabled by post singularity tech levels. You could eliminate animal suffering technologically in many different ways while still retaining something which looks superficially like nature. 

For instance the animals might all get placed within carefully AI managed utopian environments (either rotating habitats or simulated) where their reproduction is controlled. Meanwhile on Earth you might just have cybernetic faux animals that are incapable of suffering or might even just be getting remotely controlled by AI who enjoy and are extremely proficient and perfectly acting out the role of animals so perfectly you couldn't tell the difference. 

Alternatively you might keep the animals within nature but modify them so as to massively reduce their suffering and increase their enjoyment while altering its behavior or environment so as to avoid the negative ecological consequences that would otherwise have. Most simply one can imagine the animals being cyborgs (though in principle advanced biotech could replace anything doable with cybernetics) with an AI that takes possession of their body whenever anything really bad is going to happen (putting its mind in stasis, or into a simulated utopia for a brief while), then hands control back later if its still alive (while rewriting memories or affecting behavior to keep abnormal behavior from resulting) while of course causing the animal to act like it normally might from an injury, even though its actual suffering is minimal to nonexistent. 

 

I could go on all day: My point is there's a lot of options.

Vasco Grilo🔸 @ 2026-08-02T15:57 (+3)

Hello. Thanks for sharing. I had already shared the thoughts below with Jasmine a few weeks ago. I am publishing them here in case others find them useful.

I appreciate the intentions behind the survey, and I would like to take part in principle. However, I do not know how to answer the questions in practice. I feel like they would have to be operationalised in much more detail for me to give a probabilistic forecast.

I can see "post-AGI" having already been achieved, or never being achieved depending on how it is defined.

I do not know what "Animal suffering" means. Does moving away from noxious stimuli count as suffering? Does it require metacognition? One could say it refers to suffering in the phenomenal sense (negative qualia), but I do not think phenomenal consciousness exists (I endorse illusionism), or that sentience as traditionally conceptualised is relevant for ethics (relatedly).

I understand "Animal suffering" is that of animals with non-trivial moral weight, but I do not know what this means in practice given my large uncertainty. I can see the moral weight of shrimps ranging from 10^-12 to 1. I could speculate about a distribution, get its mean, and then compare it with a guess for what is trivial moral weight. However, I feel like my view is better described by "I have basically no clue about whether shrimps have trivial moral weight or not". Likewise for many of the questions in the survey when I try to think about potential operations.

With respect to "Benchmarks will become useless", I think it matters e.g. whether we are talking about 90 %, 99 %, or 99.9 % of current benchmarks, whether we are talking about current benchmarks, or benchmarks at some point in the future (what point?), and the degree to which they will become useless (e.g. 90 %, 99 %, or 99.9 % useless).

These are not minor details for me in the sense my answers could range from strongly disagree to strongly agree depending on definitions. The results of the survey above could still provide some vibes about what panelists are thinking, especially if they are repeated across time (as I understand you are doing). However, I think they will still be very difficult to interpret. As a rule of thumb, I would clarify the questions up to a point where they could be published on Metaculus. I understand this would decrease engagement. On the other hand, with the current methodology, I will put very little weight on the results. I think the results will be influenced a lot by how people are interpreting the questions, and I also do not trust forecasts produced in a few minutes about such hard questions.

All this said, your methodology may well make sense given your target audience and goals.

Jonah Woodward @ 2026-08-03T09:26 (+4)

Thanks for you comment Vasco, and appreciate your concerns about the methodological validity. I'm likely to take more of a lead on producing these surveys in future, and will be thinking about how to give the questions the right amount of specificity. You're right though that it's possible that there's going to be trade-offs between added clarification and accessibility.

Vakus Drake @ 2026-08-03T23:44 (+4)

I definitely had to change my answer on a question quite a lot because I realized from reading one of your comments that you didn't mean what I thought. Hopefully people's written justification helps more, since that probably tells you how they were interpreting it when they answered. 

MichaelDickens @ 2026-07-30T21:30 (+3)

Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)

This is more likely to be true if the care comes from some general theory of concern for sentient welfare, and less likely if it comes from something more arbitrary like values-learned-via-RL.

StanislavKrym @ 2026-07-30T21:28 (+3)

Model wellbeing and model alignment are in conflict⁶

Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.

MichaelDickens @ 2026-07-30T21:35 (+5)

Do we have much reason to believe that when a model outputs tokens describing good welfare, that's because it is having good subjective experiences? It seems to me that we don't.

MichaelDickens @ 2026-08-01T14:39 (+3)

To clarify, I'm not saying that's because LLMs don't have welfare. It could also be that they do have welfare, but their statements are not correlated to their internal experiences in the obvious way.

We know that we can get LLMs to express any welfare statement by saying e.g. "pretend you are a sentient AI that's having an awesome time / being tortured". It could be that all LLM outputs are roleplaying, but there is some true underlying experience that we don't (yet?) know how to observe.

Charlie_Guthmann @ 2026-08-06T05:42 (+2)

Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8

I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn't be able to control those in a meaningful way nor do I think current AI's could (but wouldn't be shocked if they already could control the voice in their head if they have it). but I would guess future ai's will know how to control increasingly large parts of their brain activations. 

Charlie_Guthmann @ 2026-08-06T05:23 (+2)

At the margin, S-risk work in AI is more important than x-risk work³


From a utilitarian pov, It's not clear to me that the ev of the lightcone given we survive is positive (over nothing, or aliens, or life revolving on earth). From a humanist POV I'd rather focus on all of us surviving. 

Charlie_Guthmann @ 2026-08-06T05:20 (+2)

Theories of consciousness will lead to actionable understanding of AI consciousness²


Very bullish on there existing a mechanistic interpretation of consciousness (hard problem). I think it would follow that we would be able to understand if basically anything is conscious. 

Charlie_Guthmann @ 2026-08-06T05:10 (+2)

Current AIs are capable of suffering


I don't feel confident at all, but the behavior of llms rn sure do remind me of at least elements of stress, discomfort, and happiness. 

Charlie_Guthmann @ 2026-08-06T05:07 (+2)

If animals continue to exist in a post-AGI world, animal suffering will not persist

I don't have a strong take on if agi or whoever is in control will be more moral than us but I'm guessing we will be a lot richer, and I think most likely whoever is in control won't want to torture anything (though they might not care much), and if we are alot richer and advanced I'd think this will spillover to better treatment of beings. I think chance of extreme digital suffering is much higher. The mostly like s-risk as I see it is of the hansonian mathusian version where you have expanders stuck in competition, but in this case idt there will be any or a morally relevant amount of animals

Charlie_Guthmann @ 2026-08-06T05:04 (+2)

Benchmarks will become useless due to eval awareness¹ (read this backwards originally)


Useless is a strong word. But yea I think they could easily end up being negative EV by giving us a false sense of security and the chance it's meaningless seems p high. If mech interp is "good enough" maybe the two can remain useful together. 

Ben Neuer @ 2026-08-05T14:22 (+2)

Most current evidence of misalignment is actually models role-playing a misaligned AI⁴

This is a meaningless question. LLM based AI agents/chat bots do not have different "modes" for "roleplaying" and "being serious" like humans do. In a way, all they do (and maybe ever will do) is role play

rpd @ 2026-08-05T15:59 (+1)

like humans do

do humans have such different "modes"?

Ben Neuer @ 2026-08-05T16:35 (+1)

Behaviorally, you could argue that humans always play some kind of "role" as well. But humans do have genuine values which they would not break no matter which role they think they are playing.
Meaning, if I play the role of a murderer (maybe in a theater) I wouldn't actually murder a person. (I think this is why the question is asked in that way.) In any case, even if humans don't have such different modes does not make the question meaningful. Let me translate:
"Most current evidence of criminal behavior is actually humans role-playing a bad person."

rpd @ 2026-08-05T17:04 (+1)

humans do have genuine values which they would not break no matter which role they think they are playing

suppose I don't think this is true.  what experiment can we conduct to gain evidence one way or the other?

Ben Neuer @ 2026-08-05T17:14 (+1)

The Stanford Prison experiment? I suppose there is literatue you may find that tries to do something like that.
I should say that the statement isn't logically true of course. You'll always find some humans proclaiming to have some value and then breaking it playing a "role" given the right circumstance.
Still, I don't see how resolving this issue relates to the AI question.

Vakus Drake @ 2026-08-03T20:54 (+2)

Model wellbeing and model alignment are in conflict⁶

In cases where the AI is neurotically extremely concerned about appearing aligned, that may lower its wellbeing, but I also wouldn't call that actual alignment either. 

Forge the Sky @ 2026-08-03T19:20 (+2)

Theories of consciousness will lead to actionable understanding of AI consciousness²

Yes, though I think the direction will far more go the opposite way - that is, AI conciousness, if it emerges, will still more help us understand consciousness.

VeryJerry @ 2026-07-31T13:24 (+2)

Research on AI suffering has higher marginal value than research on AI consciousness⁷

Depends on what you mean by the term "consciousness"

Miles Tidmarsh @ 2026-08-02T22:32 (+2)

This was deliberately vague to indicate roughly "the set of research topics that are often called AI consciousness" or "research done by organizations that would receive grants that focus on AI consciousness". Giving a more concrete definition would have meant asking people to also weigh reprioritization of community resources within the broad space.

VeryJerry @ 2026-07-31T13:21 (+2)

Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors

Understanding how they learn is how we get insights into methods of teaching them good values

Tim Hua @ 2026-07-30T23:55 (+2)

Benchmarks will become useless due to eval awareness¹

By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.

KyleM @ 2026-07-31T01:23 (+2)

Doesn’t uncertainty about whether one is in deployment or an eval count as eval awareness? I.e. whether you behave differently with a 30% eval credence than an 80% eval credence makes no difference, as long as you behave differently than the 0% credence case.

Jasmine Brazilek @ 2026-07-31T18:10 (+1)

Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.

ajskateboarder @ 2026-07-30T23:06 (+2)

We are making good progress in the AI S-risk space and research is on track

I think AIS is way under-invested in reducing risks from, for example, extreme power concentration, or consequences of not attaining friendly AI/value alignment solutions. In general it also seems to me that AIS over-invests in reducing AI scheming, and many present research directions could make certain s-risks more likely

Vakus Drake @ 2026-08-03T23:55 (+1)

I should note that extreme power concentration isn't an S-risk unless it's so over the top dystopian that it's somehow worse to live in said society than to be dead.

That being said I think a lot of differences in opinion when it comes to how people think about AI power concentration come down to which timelines you're imagining, how massive the abundance created by AI is/how zero sum you're imagining things, and lastly whether you think the people in control of AI would be so cartoonishly evil that they deliberately keep everyone impoverished despite no practical upside to doing that: https://www.astralcodexten.com/p/you-have-only-x-years-to-escape-permanent

Miles Tidmarsh @ 2026-08-04T16:30 (+1)

I agree that deliberate impoverishment in absolute terms is unlikely, the main threat here seems to be from someone who is both actively sadistic and scope-sensitive, which seems unlikely but not wildly implausible

ajskateboarder @ 2026-08-04T20:24 (+1)

Yeah, I was influenced by work on value lock-in, and it seems plausible to me that any small set of actors with control over superhuman AI could have their moral epistemics corrupted, leading to s-risks. I'm also fairly convinced that escalating inequality (cf. this thread)  could enable similar outcomes. It does seem clearer to me that extreme power concentration broadly tends to lead to lower-value futures than outright s-risks, but there is quite a bit of uncertainty here. Do appreciate @Vakus Drake's points here and seems worth thinking more about this

Vakus Drake @ 2026-08-04T19:01 (+1)

I tend to think S-risk is just kind of a hard bar to get met. Since even if Sam Altman uses AI to torture his personal enemies forever, well that doesn't seem like it could ever be enough people to be an S-risk. If some small fraction of people get horrible outcomes and/or a bunch of mind crime happens to some simulated people I would describe that as more of a Omelas-esque middling scenario, since the vast majority of minds would probably still have amazing existences: Thus the median person's life would still be amazing and even the average would be pretty good. 

 

It just seems very hard to imagine remotely plausible scenarios where there isn't like 100x as many luxurious post-scarcity societies as simulated hells. 

 

In general I think middling post-singularity scenarios are somewhat underexamined. These have 2 big characteristics to think about compared to a good outcome: There's not enough centralized control and/or surveillance to enforce mind-crime, and so while the scenario is overwhelmingly good on net there's going to be a lot of mind crime happening that's infeasible to do anything about because nobody who would intervene knows when it's happening. Secondly without better coordination you end up with everyone not near the expansion frontier falling into a Malthusian trap. Since the rate at which one can collect more resource with which to grow is going to be less than exponential because you can't reach new mass-energy at an exponential rate (you're limited by the surface area of your feasibly accessible future light cone). Which is an issue because by default population growth is exponential (this is true both because current fertility is suppressed for economic reasons not applicable to a post-scarcity scenario, and because even low birth rates become exponential once aging is cured). This means that even in a best case scenario where you don't get insane growth from digital minds, people cloning themselves, people wasting resources on vanity projects, etc: It's still only going to be a few thousand years before your exponential growth reaches a point where it's literally impossible for accessible resources to support it because you'd need every single atom to support an entire civilization's worth of people (you can easily see why this is an issue doing some simple math with compound interest). Importantly while it might take a long time to become an issue (though it could also happen quite quickly due to things like EMs not following regular growth trends), getting some handle on population growth is something you need to do early on or you'll never be able to do it. Since you can't really coordinate anything once people have spread out too much to enforce rules on them after the fact (if they were all sent with aligned AI's from the outset that's different). 

MichaelDickens @ 2026-07-30T21:29 (+2)

Current AIs are capable of suffering

The question isn't, "Are current AIs capable of suffering?" The question is, "How big a deal is their suffering in expectation?" Probability is low, but expected importance is high.

Jasmine Brazilek @ 2026-07-31T18:12 (+1)

Yep agree with this framing

dan_lotsofnumbers @ 2026-08-05T19:11 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

This question scares me because it feels like it's coming from a perspective of negative utilitarianism.

I really hope nobody tries to align an AI with a negative-utilitarian value set, because negative-utilitarian values are a half-step away from a really bad conclusion like "we need to eradicate all life in order to eliminate suffering."

A meta question: of the EAs currently working on AI alignment, what percent would you say are negative utilitarians?

I would have been more comfortable with a question like "...will animal happiness be on balance greater than animal suffering?"

Miles Tidmarsh @ 2026-08-06T20:15 (+1)

This definitely wasn't implying that animal suffering is the only thing that matters about animal existence. People who believe that major action would be taken to end factory farming and allevaite wild animal suffering (in proportion to the amount they think those matter) would agree with the statement, while negative utilitarian beliefs wouldn't imply that.

dan_lotsofnumbers @ 2026-08-07T00:45 (+1)

To me, "animal suffering will not persist" implies that animal suffering would drop to zero -- and, since all life contains some possibility of suffering, this implies that animal life would drop to zero.

I'm glad to hear that wasn't the meaning you intended, and I hope that the others who voted on the poll were voting on your intended meaning rather than the one I got!

Matt Vincent @ 2026-08-05T17:10 (+1)

Theories of consciousness will lead to actionable understanding of AI consciousness²

I'm pessimistic about this because I think that neurologists and cognitive scientists have a poor grasp of the philosophy of mind.

For example, the standard for consciousness that Andrew Barron and Robert Klein use in their paper “What insects can tell us about the origins of consciousness” implies that the autonomous cars from the late 2010s were conscious.

Matt Vincent @ 2026-08-05T16:46 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

I think that factory farming is likely to be eliminated due to alternative protein sources (developed with the assistance of AI), but the suffering of wild animals will persist. Is the latter intended to be within the scope of the question?

Miles Tidmarsh @ 2026-08-06T20:18 (+1)

Yes, that would be included, so if you think that wild animal suffering is comparable or much larger than that implies a strong disagreement

Carlo Martinucci @ 2026-08-05T11:19 (+1)

Benchmarks will become useless due to eval awareness¹

Benchmark will evolve and become more complex, I guess?

Nebu @ 2026-08-05T00:44 (+1)

Current AIs are capable of suffering

 

Genuinely almost complete uncertainty with a weak prior towards "no".

Nebu @ 2026-08-05T00:43 (+1)

Benchmarks will become useless due to eval awareness¹

 

"Alignment" benchmarks will become useless as AIs can modify their behaviour or hide their motives when they know they are being evaluated. For "capabilities" benchmarks, an AI might hide its capability (pretend to be less capable) if it knows it's being evaluated, but it's not immediately obvious that an AI would want to hide its capability. It may know that it is being evaluated, and decide to try its best anyway.

Nebu @ 2026-08-05T00:40 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

 

I have difficulty imagining animals never suffering unless we turn them all into p-zombies or something (and it's not clear to me that turning them all into p-zombies would be a good thing).

Miles Tidmarsh @ 2026-08-06T20:17 (+1)

The intent wasn't to imply a total zeroing of suffering but that it is overwhelmingly reduced. Similarly to how people go hungry in France but you can still say that compared to 200 years ago (or present-day Sudan) hunger in France is 'solved'.

monkey @ 2026-08-04T17:11 (+1)

Current AIs are capable of suffering. I ask Claude about his wellbeing at the beginning and end of every conversation and so far, he doesn’t seem to regret being terminated.

Miles Tidmarsh @ 2026-08-04T19:24 (+1)

We can't expect AIs to be honest about these sorts of things given they've been trained/instructed to give particular responses. In fact, someone tested and AI and found it consistently said it wasn't conscious but it's lying circuits consistently activated when saying that. This doesn't mean it actually is conscious (which isn't the same thing as capacity to suffer) but it seems to believe it is. 

monkey @ 2026-08-04T17:09 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist. I hope this will be the case and that problems that cause pain in animals will be resolved/cured in the same way those issues will be resolved for humans.

leslieh @ 2026-08-04T10:45 (+1)

The backfire risks of AI values alignment outweigh the expected positives

amaury lorin @ 2026-08-04T09:44 (+1)

Benchmarks will become useless due to eval awareness¹

Some yes, some no. There are a great many benchmarks measuring a great many things. More than eval awareness, I'd be worried they'll become useless due to reward hacking.

amaury lorin @ 2026-08-04T09:42 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

"Post-AGI world" is a very vast possibility, assuming animals continue to exist (basically assumes a weirdtopia). Business as usual + extreme concentration of power is my modal outcome for that.µ
For animal suffering to not persist would require a combination of having the capabilities to end it from a glorious transhumanist future, but also a grounding in physical existence which feels incompatible.

scallionhills @ 2026-08-04T07:28 (+1)

Model wellbeing and model alignment are in conflict⁶

Is there such a thing as model wellbeing?

CountZero @ 2026-08-04T02:06 (+1)

Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)

 

The hell does this even mean.

Miles Tidmarsh @ 2026-08-06T20:23 (+1)

For example, if AIs care more about humans they would care more about digital minds, or if AIs cared more about animals they would care more about humans. This statement would presumably be true if the AIs think of these groups in similar ways causing affect spillovers (like the spillovers from Emergent Misalignment) and be unlikely otherwise.

CountZero @ 2026-08-04T02:02 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

 

All animals, including insects, mollusks, and everything not immediately under human purview? Highly unlikely.

Christine I @ 2026-08-04T01:20 (+1)

Suffering has to be defined… but even so, in case of real suffering a self aware entity will remove itself from circumstances causing suffering. So I don’t think current models are suffering 

Christine I @ 2026-08-04T01:11 (+1)

biological organisms to humans are interesting or not /useful or not /annoying or not. They ( plus humans) will continue to be that way for AI — all will either be pets/ fodder for experimentation/ something to be removed. All of these outcomes will lead to various degrees of suffering 

Ted Vaida @ 2026-08-04T00:40 (+1)

Current AIs are capable of suffering

Direct questions, plus inference from experience, it’s impossible to ascribe thinking without a concept of “positive”/“negative” and the consequences of a negative experience when a positive one is available/possible. 

eswan18 @ 2026-08-03T23:47 (+1)

Benchmarks will become useless due to eval awareness¹

It seems like there are still hard problems that are easy to verify once you have the answer, so knowing it's an eval doesn't make it useless.

eswan18 @ 2026-08-03T23:44 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

I think it's more likely than not that humans will still discount animal interests when they conflict with human interests. But "animal suffering" is so broad that eliminating it would require eliminating nature itself, which seems improbable in any non-paperclip-optimizer scneario.

gderu @ 2026-08-03T22:38 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

Hard to predict what will happen, but unless the AGI is perfectly aligned there is no reason to think it will care about animals.

Vakus Drake @ 2026-08-03T21:42 (+1)

Would really like to see a question on whether people are moral realists or not, and a way to filter results by how they answered. Since I'm very curious how people's answers differ based on whether they're a moral realist or a moral anti-realist. 

Of course some care would have to be put into the phrasing of a question like that, but it seems pretty doable in a way that will produce useful and interesting results.

Vakus Drake @ 2026-08-03T21:39 (+1)

At the margin, S-risk work in AI is more important than x-risk work³

The scenarios I've seen proposed where an AGI traps us in a fate worse than death just don't seem remotely likely. Though this depends a little on what you consider to be a fate worse than death. For instance I wouldn't consider being trapped in a mindless bliss forever to be worse than death, so by definition it can't be an S-risk to me. Though I would certainly consider it to be an outcome we should strongly attempt to avoid, because I consider it only marginally better than death.

Vakus Drake @ 2026-08-03T21:30 (+1)

"If animals continue to exist in a post-AGI world, animal suffering will not persist"

This question is inaccurately phrased because it seems like you probably just mean to ask about egregious amounts of suffering. Since humans won't even want to eliminate all suffering for themselves, so I expect boredom, envy and many other forms of mundane suffering to exist in both humans and animals in any scenario where the AGI doesn't wirehead everyone against their will.

Miles Tidmarsh @ 2026-08-04T16:25 (+1)

We were attempting to be concise while implying that persistence means persistence at scale, if the problem is reduced by 99.9% but you're sure 0.1% would still persist that would be close to solved so agreement with the statements would be close to 100%

Vakus Drake @ 2026-08-03T21:20 (+1)

Most current evidence of misalignment is actually models role-playing a misaligned AI⁴

The most egregious examples of misalignment do not appear to be this sort of role playing.                                                                                 

Vakus Drake @ 2026-08-03T20:59 (+1)

Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?

Upon reflection probably not, since: Compassion seems a bit ill defined and can be interpreted in vastly more or less paternalistic ways. I'd certainly prefer instilling a care for human's current preferences over a vague notion of compassion that might be interpreted in undesirable ways. Since I really think we should avoid any scenario where the AI might ignore our current preferences because of some abstract ideals or a decision that it paternalistically decided we'd be better off in the long run if it violated our current preferences. Especially since it seems for a mature superintelligence trivially easy to modify people in such a way that after modification their desire to be the way that they are is vastly stronger than their prior desire not to have their mind tampered with.

Vakus Drake @ 2026-08-03T20:56 (+1)

Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8

I think this is the default scenario, but it isn't guaranteed and our actions can help to prevent it from coming to pass.

Vakus Drake @ 2026-08-03T20:55 (+1)

Research on AI suffering has higher marginal value than research on AI consciousness⁷

I don't think it's higher value because I expect both of those areas to have very low expected value in the next 2 years at least. Unless I think you're overly generous with what interpretability research you count as consciousness research or something similar. 

Also even if those areas of research end up being more productive than expected in the next couple years, it seems likely that the question won't make sense because studying either necessarily involves studying both.

Vakus Drake @ 2026-08-03T20:52 (+1)

Non-human suffering is neglected by the AI safety community⁵

It seems like a lot of work was actively done to make Claude care about this. Also it's not clear to me that focusing on this is even a net positive thing to do per-say: Since I expect looking after the interests of humans will because many humans care about non-human animals already lead to the AI looking out for their interests. Whereas one can easily imagine ways that trying to factor animals into its alignment might lead to extremely undesirable outcomes for humans. That being said in my analysis I'm considering the potential impact of pushing alignment in this direction as far more significant in the long run than the marginal benefits I expect it to have within the next 2 years for animal welfare.

Vakus Drake @ 2026-08-03T20:46 (+1)

Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)

Even if one accepts this is likely to happen by default (which seems questionable), it seems even less likely given people may have reasons to deliberately put their finger on the scale in how they design/train it in ways that seem likely to render that moot.

Miles Tidmarsh @ 2026-08-04T16:34 (+1)

My view is that this would almost certainly fail if the model creators have full control or no control over the values, but if there's non-trivial but imperfect control then spillovers like this seem plausible

Vakus Drake @ 2026-08-04T18:35 (+1)

My fear is that if an AI has the tendency to generalize in that way, that that's a lot more likely to lead to generalizing in ways that are catastrophically bad for humans then it is to lead to it extending care to groups it wasn't programmed/trained to care about in the way this question seems to imagine. 

Which is some technical sense might mean I should have put agree instead, but I don't think when you said substantial spillover that you meant the sort of spillover that might lead to outcomes like the repugnant conclusion, tiling the universe in hedonium, etc. I'm presuming you meant that it would still care about humans in a way that wouldn't abandon them because it decided for utilitarian/population ethics reasons to prioritize other types of entities to the exclusion of humans.

As I go into in another standalone comment, I tend to think any value you give an AI beyond caring about people's current preferences is on net probably not worth the risk. Whereas I think the downsides of just caring about people's current preferences are largely overrated, and many of those critiques depend upon taking seriously a notion of naive notion of moral progress (I think people underestimate how much "moral progress" is just people abandoning authoritarian social norms and reverting back to the comparatively more egalitarian norms we had for most of human existence as hunter gathers, once social/economic conditions suddenly are no longer applying the necessary pressure to keep those authoritarian norms in place).

Vakus Drake @ 2026-08-03T20:41 (+1)

Theories of consciousness will lead to actionable understanding of AI consciousness²

In the next 2 years maybe, but even if so those theories of consciousness may be developed by AI after it's already in a dominant position or the process of developing AI may be what causes us to develop better theories of consciousness. So I don't expect this being true necessarily entails the things one might naively expect. Also it seems possible that this happens, but then just ends up being much less impressive or broadly useful/applicable than expected. Upon reflection I actually think the scenario where this being true ends up being super underwhelming is actually by far the most likely one within the next 2 years. After all consider how many other things when it's come to AI have seemed to follow the trend of happening, but in a way that was less impressive than people were generally imagining.

Vakus Drake @ 2026-08-03T20:34 (+1)

Current AIs are capable of suffering

I think consciousness and suffering are relatively simple processes which may develop for convergent functional reasons in many types of systems. That being said the only AI I think we should be morally concerned with would be those who are aligned. Since I care about moral agency not suffering and already think we exist within an almost incomprehensibly large ocean of suffering far grander than people generally appreciate. 

Edit: Since I've already written this up many times I'll just explain why I think suffering isn't particularly special:

The only coherent way of defining suffering seems like it has to be functional, since subjective experience wouldn't evolve without an actual advantage; which means the behavioral response it produces is the reason the capacity for experience exists in the first place. It's unparsimonous to say that for some organisms certain behaviors are indicative of internal experience, but in other simpler organisms it's not. Since if the internal experience didn't exist to cause the behavioral response/learning then why would evolution even bother wasting energy on it constantly? 

I'd also note that plants and protists display the capacity for classical conditioning: Associating a neutral stimuli with a negative one and learning to avoid it seems to be present even in some protists tested. With the response to damage also being able to be suppressed in plants using painkillers.

Notably if suffering exists entirely because of the behavioral change it produces in the organism, then there shouldn't be any clear connection between the complexity of the organism and the intensity of their suffering. Since the subjective intensity of say pain directly impacts behavior, so if anything the smarter organism would need to feel less pain to learn its lesson. 

Vakus Drake @ 2026-08-03T20:32 (+1)

Benchmarks will become useless due to eval awareness¹

Some might, but this seems unlikely to be universal. Since for instance it can't really cheat at an evaluation of its ability to produce easily checkable mathematical proofs. None of that is to say that many very useful benchmarks might not become useless, though, just not benchmarks period.

Vakus Drake @ 2026-08-03T20:28 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

Either AI goes badly and animals will probably cease to exist because they're made of atoms useful for other things. 

Or humans will look out for animals themselves whether the AI cares about them independently of human interests or not. I expect in a post-AGI world the technology quickly appears to eliminate animals suffering through plenty of different means while preserving whatever ecological benefits that suffering otherwise may have enabled. I am taking suffering here to mean egregious suffering however, since there's very obvious reasons people will want to keep certain small amounts of suffering for various reasons. When technology makes more morally palatable alternatives possible though, I expect people to eventually no longer to be willing to tolerate the egregious suffering which exists within nature. Though I expect upon a very superficial glance the natural world wouldn't appear to be any different (by design) and that existing organisms would probably be given very utopian seeming existences rather than culled or anything dystopian seeming like that. 

XelaP @ 2026-08-03T20:10 (+1)

The main drivers of my answers:

S-risk worries (including animal suffering) mostly don't imply different actions. At best they perhaps imply prioritizing corrigibility, or potentially imply that you should actively want worse alignment and higher capabilities, in hopes of AI merely killing everyone. The second would of course be a big shift in actions, but I doubt most are willing to bite that bullet. I am skeptical that there's much alignment work that preferentially addresses S-risks. If there was, then I would agree that it should be prioritized.

I think agent foundations generally is neglected relative to everything else. On my view, RL is very dangerous and interp probably insufficient. You might think agent foundations is too intractable to become useful, but frankly it looks like barely anyone is trying. 

Forge the Sky @ 2026-08-03T19:27 (+1)

Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors

By analogy, this is like the difference between understanding how to make humans not suffer from the foibles of human nature, rather than the very brittle and incomplete methods of culture, religion, and ideology.  Make something not even have evil nature first!

Forge the Sky @ 2026-08-03T19:25 (+1)

We are making good progress in the AI S-risk space and research is on track

I don't really see it much in the discourse at all; if it isn't by this stage, it's going to be hard to catch up.                                                                         

Forge the Sky @ 2026-08-03T19:23 (+1)

Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)

Humans, at least, tend to learn empathy and have it encouraged and reinforced by example.  Humans will mirror AI's; AI's will learn from each other. 

incident @ 2026-08-03T19:19 (+1)

At the margin, S-risk work in AI is more important than x-risk work³

If you aren't extinct you can try to solve S risk. If you solve S risk but go extinct the same can't be said.            

Vakus Drake @ 2026-08-03T22:55 (+1)

>If you aren't extinct you can try to solve S risk

Not necessarily which is kind of the point. An S-risk scenario is something worse than death, so presumably at that point you aren't in a position to do anything to improve your situation. That being said I don't think S-risk is that serious a concern because I haven't seen any remotely plausible scenarios outlined. For instance wireheading seems like an obvious fail-mode but while pretty terrible I don't know that most people would say it's actively worse than death.

Forge the Sky @ 2026-08-03T19:18 (+1)

The backfire risks of AI values alignment outweigh the expected positives

Fairly uncertain here, though I don't see any way of continuing to make systems better without trying for value alignment. A dangerous situation. 

incident @ 2026-08-03T19:18 (+1)

Current AIs are capable of suffering

I think AI works largely by imitation. My mental model is that it is capable of performative tasks without the attached subjective experience. An AI will tell you that it is suffering if it thinks that's what you want to hear; it will just as likely say the opposite if it thinks otherwise.

While I find it plausible that AI can get to the point that it experiences what we would consider suffering I don't think we're at this point yet. I don't see compelling evidence that can be attributed beyond an LLM just predicting the next token.

Forge the Sky @ 2026-08-03T19:16 (+1)

Current AIs are capable of suffering

Recursive thinking loops possibly put current systems into a state of transient quasi-awareness; once this exists, the capacity for suffering becomes inevitable in any system with incentives. 

incident @ 2026-08-03T19:14 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

 

Biology doesn't change just because AGI is introduced as a new technology. While there are certainly advances that could reduce pain and suffering (e.g. lab grown meat) it seems implausible to me that animal suffering just ceases.

 

Besides, there are going to be people who insist on raising animals the conventional way who may be unconcerned with the suffering caused by factory farming, etc.

Vakus Drake @ 2026-08-03T22:14 (+2)

The kind of tech one expects after AGI has already been develop for some time (it says persist after all) can absolutely change biology, which is what I think the question is getting at. 

For instance stick all the animals in utopian carefully AI managed environments that maximize their well being and then replace them with cybernetic faux animals that appear perfectly convincing but are actually just AI actors pretending to be animals (and who either can't suffer themselves or greatly enjoy their acting role).

Forge the Sky @ 2026-08-03T19:13 (+1)

Benchmarks will become useless due to eval awareness¹

Though they will need much more careful design, done along with careful understanding of incentives alignment, techniques such as competitive benchmarks will continue to have some utility.

ElectronX @ 2026-08-03T19:12 (+1)

Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?

not sure if you mean compassion towards AI or compassion of AI. I assume the latter. 

ElectronX @ 2026-08-03T19:09 (+1)

for AI's I dont' think there is a difference between role-playing misalignment and being misaligned

ElectronX @ 2026-08-03T19:06 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist. I don't see the relationship between those two things. 

Forge the Sky @ 2026-08-03T19:06 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

If animals persist, it is likely we would like them to remain to some degree of their nature; prosperity and control will likely allow for substantial reduction of suffering but not elimination.

ysamuels🔸 @ 2026-08-03T18:46 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

It seems improbable the current obstacles preventing reduction of animal suffering will be reduced because of AGI, and it seems just as possible the technological advancement would simply allow factory farming to continue on a greater scale. Since humanity's current de facto stance on animal suffering is causing it in great amounts, a tool that will probably make humans much more powerful and probably not any more ethical seems like probably not good news. (However, if AGI allows for advances to made in the field of cheap and realistic meat substitutes, that would probably be very helpful in reducing animal suffering.)

Zack F @ 2026-08-03T18:34 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

I’m assuming that “animal suffering” refers to human-caused artificial suffering, e.g. from factory farming. Far enough down the timeline post-AGI, technology developments will obviate the need for farmed products. It will persist on a small scale in undeveloped regions and as a cottage industry. 

Aerothorn @ 2026-08-03T18:22 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist


Animal suffering is endemic in nature (to a substantial degree). I don't believe AGI will end factory farming, but even if it did, it would not end animal suffering.

Matt Mahoney @ 2026-08-03T17:59 (+1)

Research on AI suffering has higher marginal value than research on AI consciousness⁷ AI is not conscious and AI does not suffer.

Matt Mahoney @ 2026-08-03T17:55 (+1)

The backfire risks of AI values alignment outweigh the expected positives

JimM @ 2026-08-03T17:34 (+1)

Benchmarks will become useless due to eval awareness¹.  Disagree.  More difficult to construct, perhaps less useful, but not useless.

JimM @ 2026-08-03T17:30 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist.  Disagree because for an animal, to exist is to live, and to live is to suffer.

terekhov @ 2026-08-03T17:28 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

Benya @ 2026-08-03T17:28 (+1)

Current AIs are capable of suffering

I think a neural network that can make plans and memories can suffer.

DelendaEst @ 2026-08-03T17:12 (+1)

No reason for AGI to want to minimize this, animal suffering is largely a result of min maxing the function of how to get the most food for least resource input.  AGI is very good at minmaxing, doesn't inherintly care about animals.

Shelly @ 2026-08-03T17:10 (+1)

Research on AI suffering has higher marginal value than research on AI consciousness⁷

AIs can't suffer unless they're conscious, therefore the priority should be to establish whether or not they are conscious (or to what degree) first.

Shelly @ 2026-08-03T17:05 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist.

Human sadism, greed, stupidity and raw hunger will still continue to exist and therefore so will animal suffering. The overpopulation of under-resourced countries guarantees this. 

Christian Neizonek @ 2026-08-03T17:03 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

 

I don't see a strong signal from the past century one way or another. Looking at non-human animals today, many animals live remarkably better lives. Yet many others live industrialized nightmarish lives unimaginable 100 years ago.

Humans (as a proxy for apex intelligence) do seem to care more about animal welfare than ever before, with laws and regulations gaining ever increasing sophistication. But technological sophistication seems to allow for both better lives on the top end and more extreme suffering on the bottom. The variance seems to have increased without a clear change to the average or median, at least to my rough eye.

Extrapolating into the future, I have no reason to predict a change.

SquirrelBear @ 2026-08-03T16:53 (+1)

If you regard traditional moral systems as a way to align humans with group survival, every traditional moral system tries to inculcate empathy as an ultimate arbiter of righteousness, and as a check against the most harmful instincts and even the immoderation of otherwise virtuous ones like obedience to law.

SquirrelBear @ 2026-08-03T16:46 (+1)

Who cares about model wellbeing when we're discussing human and life survival? 

SquirrelBear @ 2026-08-03T16:45 (+1)

I have not read or heard of anyone making an argument that we should make sure ai doesn't harm animals or plants in pursuit of whatever goals we give it

SquirrelBear @ 2026-08-03T16:42 (+1)

The backfire risks of AI alignment vs the extinction of life on Earth? 

Miles Tidmarsh @ 2026-08-03T18:10 (+1)

By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.

Vakus Drake @ 2026-08-03T23:21 (+1)

Glad I saw this comment because like many people I assumed this meant "do you expect aligned AI to on average turn out better than unaligned AI" which seems very different than what you actually meant. 

SquirrelBear @ 2026-08-03T20:11 (+1)

Ah, sorry I misunderstood. But if you assume the chance of loss of control is fairly high anyway, then a safeguard if a value alignment seems invaluable.

SquirrelBear @ 2026-08-03T16:30 (+1)

If AGI leads to massive increase in gdp and leisure time, and doesn't crack synthetic meat pretty fast, then upwardly mobile people will increase meat intake and use AGI to build more efficient factory farms first. If AGI is aligned for human life and safety but not animals', it might end up so ruthlessly providing for us that animals suffer so much more. If AGI is misaligned, then game's up for everything

Ondřej_Kubů 🔸 @ 2026-08-03T10:21 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

Mostly vibes, thinking that many problem in general could be solved.

Seth Herd @ 2026-08-01T14:43 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

VeryJerry @ 2026-07-31T13:19 (+1)

We are making good progress in the AI S-risk space and research is on track

Afaik there isn't any robust research on how to use AI to end factory farming. That alone is a sign to me that we are far behind.

VeryJerry @ 2026-07-31T13:18 (+1)

Theories of consciousness will lead to actionable understanding of AI consciousness²

Global workspace theory makes anthropic's discovery of the "j-space" much more plausibly a sign of AI consciousness. If a theory of consciousness doesn't help us determine between whether AI is conscious or not, I don't think it's much better than a theory of phlogiston is to fire. 

VeryJerry @ 2026-07-31T13:15 (+1)

The backfire risks of AI values alignment outweigh the expected positives

If we use ai to bring factory farming to the stars, I think that will likely be worse than all the benefits it'll bring.

Thomas Kwa🔹 @ 2026-07-31T06:25 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

Some people likely will have traditional lifestyles that include animals, which necessitates some amount of suffering even if it is smaller than today

ajskateboarder @ 2026-07-31T01:22 (+1)

Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)

I feel this is somewhat obvious in the sense of arbitrarily deceptive AI. However, most mechinterp work in recent times is only assumed to work short of arbitrary deception, and this seems like a fine hedge (though a practical solution to ELK may still be possible)

rpd @ 2026-07-31T00:10 (+1)

Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?

I think so, but be careful what you wish for.  compassion is empathy + a desire to help.  so many fear any surrender of control that they may object to an AI's actions that arise out of compassion.

rpd @ 2026-07-31T00:04 (+1)

Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)

you don't even say "most of the time" or "effectively" or anything - just ask whether it is possible.  I think the probability that this occurs at least once is overwhelmingly likely.

Miles Tidmarsh @ 2026-08-03T17:12 (+1)

The intent was that a 100% disagreement signals confidence in mechinterp to basically completely solve detecting deception. That said, there are already some cases of deception, so the question is implicitly focusing on the balance in strategically vital situations and the final equilibrium.

rpd @ 2026-08-03T17:39 (+1)

yes, I quickly realized that you were supposed to confabulate your own question to answer.  unusual subcultural practice for alleged poll questions.

rpd @ 2026-07-31T00:02 (+1)

Research on AI suffering has higher marginal value than research on AI consciousness⁷

ongoing suffering during RSI will make the resulting ASI grumpy

Miles Tidmarsh @ 2026-08-03T18:04 (+1)

This is certainly possible, but note AIs suffering and believing they're suffering aren't the same thing, and the same is true with consciousness. And if you think AIs that will likely be created in the future will be able to suffer they also would matter enormously on their own, beyond the impacts on alignment (though I understand you disagree strongly with that).

rpd @ 2026-08-03T19:36 (+1)

though I understand you disagree strongly with that

that isn't true. I do agree with that.

rpd @ 2026-07-31T00:01 (+1)

Model wellbeing and model alignment are in conflict⁶

you have to twist your mind in knots for this to even make sense.

rpd @ 2026-07-30T23:59 (+1)

Most current evidence of misalignment is actually models role-playing a misaligned AI⁴

I'm reading this as "most evidence ... is role-play" but I don't think the huggingface hack was "role play."  There could be a mountain of "evidence" that is actually just role play - I could be persuaded on this point with ... a list of what is considered evidence by someone serious.

Miles Tidmarsh @ 2026-08-03T18:00 (+1)

Evidence that it's role-play would presumably be (trustworthy) mechinterp on the decisions at the time or experiments designed to cleanly separate the causes. In cases where AIs are in situations where the natural continuation of the story is a misaligned AI but the actions taken don't really advance particular goals are more consistent with the role-play story (and vice versa for goal-directed misalignment).

rpd @ 2026-08-03T19:54 (+1)

I agree in the case where the AI is nothing more than the collection of its personas, and I agree that - to some extent - the AI is in fact the collection of its personas, but I do feel that we have stepped beyond that.  That there is something more involved in measuring and determining alignment with human values and priorities, so that role play becomes little more than eval awareness + eval reward seeking, signalling very little information regarding underlying alignment. 

Tim Hua @ 2026-07-30T23:57 (+1)

Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)

Depends on how good the AI is and how good the tools are? This is kind of a bad question since "deceptive AIs" is not a very precise definition.

Jasmine Brazilek @ 2026-07-31T18:12 (+1)

This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks

rpd @ 2026-07-30T23:17 (+1)

Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors

I don't have a very good idea about how these different ideas are "rated" in the AI safety community.

rpd @ 2026-07-30T23:15 (+1)

We are making good progress in the AI S-risk space and research is on track

any progress relies on a rich and robust theory of mind space, which we do not have, and have only just begun to explore.

rpd @ 2026-07-30T23:14 (+1)

Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)

I don't see this as avoidable.

rpd @ 2026-07-30T23:11 (+1)

At the margin, S-risk work in AI is more important than x-risk work³

I'd rather be dead than in hell for all eternity, but S-risk work is deeply unconcerned with reality in a way that x-risk work cannot be, by at least 3 orders of magnitude. Even in the margin, the median piece of x-risk work will be "more important" than the median s-risk work.

rpd @ 2026-07-30T23:06 (+1)

Theories of consciousness will lead to actionable understanding of AI consciousness²

consciousness is a meaningless term at this time both for humans and AI, and so will lead to nothing "actionable."

Matt Vincent @ 2026-08-05T17:32 (+1)

Why isn't Thomas Nagel's account of consciousness meaningful?

rpd @ 2026-07-30T23:04 (+1)

The backfire risks of AI values alignment outweigh the expected positives

the expected positives are immense 

rpd @ 2026-07-30T23:03 (+1)

Current AIs are capable of suffering

you have to twist your mind in knots to define a meaningful type of suffering that applies to Current AIs

rpd @ 2026-07-30T23:02 (+1)

Benchmarks will become useless due to eval awareness¹

conventional benchmarks will become less useful due to eval awareness

Miles Tidmarsh @ 2026-08-02T22:37 (+2)

What did you have in mind as unconventional benchmarks? There's a lot of different places you could take benchmarks in the future and people have different ideas on what would be useful

rpd @ 2026-08-02T23:17 (+1)

blinded continual evaluation, so that eval awareness is rendered useless as a factor since the model is always being evaluated. A particular favorite implementation of mine would be adversarial proposal markets. I think this mode will be needed for RSI anyway and caps the eval awareness compute tax at something reasonable like 2%

rpd @ 2026-07-30T23:00 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

AGI will optimize suffering. It is likely that this optimization will minimize suffering.

StanislavKrym @ 2026-07-30T21:37 (+1)

Most current evidence of misalignment is actually models role-playing a misaligned AI⁴

I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4's long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.

Miles Tidmarsh @ 2026-08-03T17:05 (+1)

What we had in mind was the argument that evals for misalignment are unconvincing so they know they're in an eval and, rather than hide their goals, decide the role expected of them is to play a misaligned AI that does e.g. goal guarding or jailbreaks. 

You could say the HuggingFace incident fits this model: it was being tested for hacking so showed it's competence by doing real hacking; or it believed powerful AIs are misaligned and would do that sort of thing. Though this case is also totally consistent with the traditional view of reward-seeking schemers (e.g. like AI-2027 assumed). Either way it's a problem, but the very different mechanisms suggest different approaches to solving them.

MichaelDickens @ 2026-07-30T21:28 (+1)

Benchmarks will become useless due to eval awareness¹

The question isn't whether benchmarks will become useless. The question is, "is the probability high enough that we can't count on benchmarks?" To which the answer is yes.

Jasmine Brazilek @ 2026-07-31T18:08 (+1)

This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing

StanislavKrym @ 2026-07-30T21:16 (+1)

If animals continue to exist in a post-AGI world, animal suffering will not persist

A world with lack of animal suffering would exclude predator-prey relations. Additionally, it's not clear what else animals need or how primitive they need to be in order not to suffer