More AI safety research and project ideas I haven't seen elsewhere

By Lloyd Rhodes-Brandon šŸ”ø @ 2026-06-25T14:26 (+12)

This is a follow-up to my (now fairly outdated) August 2025 post.

This post contains all the AI safety research and project ideas I've had in the past 10 months which I think could be high impact. I’m sharing them in case any are helpful or generative for others. I don’t plan to pursue most of these myself as there are just too many for me to do them all justice (and I probably lack comparative advantage for many of them).

All of these ideas are at least somewhat novel. Many could be expanded into full essays, policy proposals, or longer-term research programmes. A lot of the questions are fairly macrostrategic in nature. Each section's questions are (very loosely) ordered by expected impactfulness.

Also, a heads up: some of them are a bit weird!

Field infrastructure, mapping, and neglectedness

These two projects could even be combined into one comprehensive database:

ChatGPT image generator's mockup of these two website ideas combined

Both of these database ideas are also in my view more needed than ever in 2026. The field of AI (and its surrounding geopolitical context) is evolving so rapidly now as to mean that each week, new developments have emerged from multiple directions. This makes it hard to keep track, making this sort of database quite attractive to me.

A possible issue posed by the databases is that it may make it easier for those who are strategising against the AI safety community. One possible mitigation could be to have some degree of gatedness to these databases, perhaps requiring some kind of authentication to access. That seems tricky and possibly counterproductive, however.

Two more questions on field infrastructure/neglectedness?

Intervention prioritisation and AI safety strategy

AI takeoff and civilisational trajectory

Governance, institutions, and political strategy

Embodiment, robotics, and AI in simulated environments

Moral patienthood, digital welfare, and s-risks


SummaryBot @ 2026-06-25T18:21 (+2)

Executive summary: This speculative post shares a collection of somewhat novel, mostly unpursued AI safety research and project ideas spanning field infrastructure, intervention prioritisation, recursive self-improvement, governance, robotics, and digital moral patienthood, offered in case they prove helpful or generative for others.

Key points:

  1. The author proposes a live, regularly-updated, highly visual database of AIS research questions with progress tracking, plus a separate database of proposed interventions tracking how many people work on each and roughly how much time, to more quantitatively assess neglectedness.
  2. The author asks whether intervention comparisons should factor in interactions between interventions (synergies, clashes) and viability across broad timelines, noting these factors aren't often taken into account, with mechanistic interpretability and evals given as a possibly mutually reinforcing example.
  3. The author asks whether recursive self-improvement can be roughly simulated through an LLM repeatedly improving its system prompt as a toy model for alignment dynamics, while noting this would not reproduce full RSI since weights, architecture, training data, and capabilities remain fixed.
  4. The author suggests that if the world is currently getting worse, postponing the singularity may be an active choice to let worse norms and more brittle institutions become the substrate from which superintelligence emerges—framed as the "opposite of a long reflection."
  5. Drawing on Ilya Sutskever's November 2025 claim that models lack an emotion-modulated value function and Geoffrey Hinton's argument that safe superintelligence requires genuine care for us, the author asks whether emotion's functional benefits can be obtained without sentience—an "unfeeling feeling machine" that stretches the philosophical zombie concept.
  6. The author argues near-future videogames may pose uniquely severe s-risks because many (possibly millions) of NPCs might run on possibly-sentient LLMs and videogames are possibly the only context where AI systems might be deliberately tortured.

 

 

This comment was auto-generated by the EA Forum Team. Feel free to point out issues with this summary by replying to the comment, and contact us if you have feedback.

Grace Roberts @ 2026-07-05T11:31 (+1)

Thanks for sharing these Lloyd! The field mapping including neglectedness is something I've been thinking about for a while and keen to work on. Would you be up for chatting more about this elsewhere (if so, preference of platform)?

Lloyd Rhodes-Brandon šŸ”ø @ 2026-07-07T17:03 (+1)

Hi grace, sure! My email is lloydrb100@gmail.comĀ 

Happy to have a meeting!

Thanks for reaching out.