Total Safety Transparency?

By Austin @ 2026-09-22T05:43 (+20)

This is a linkpost to https://manifund.substack.com/p/total-safety-transparency

The AI safety movement should push itself to be dramatically more transparent to the public.

To date, the AI safety movement has been one of the strongest forces for clarity and wisdom in the world. The movement has been prescient on the subject of concerns from existential risk, seriously grappling with outcomes others dismissed as sci-fi nonsense.

Society is now waking up to the potential threats of advanced AI. I understand that many in the movement are feeling the crunch, and thinking more carefully about optics and what they publish. Even so, acting transparently is more important than ever.

Why transparency?

If you want labs to be transparent, you should be transparent too. AI 2040 proposes “Total Research Transparency” for labs to open up their research, algorithms, LLM weights. Safety should do likewise. Model good behavior, to convince labs that this is an acceptable and correct way to behave.

I think AI safety people are unusually virtuous; you should display that virtue. “Nor do they light a lamp and then put it under a bushel basket; it is set on a lampstand, where it gives light to all in the house.” (Matthew 5:15)

Transparency ties you to the mast, forces you to be virtuous. Famous maxim: “Act as though what you do might end up on the front page of the NYT”. And, what better way to enforce that than to publish everything you think and do?

You can’t keep things private anyways, given stylometry and cheap intelligence. Actions cast a shadow in the world, and AI will be able to detect that shadow, reconstruct that action. (More here.)

Transparency was a cornerstone of this movement. It is part of what drew me (and many others) to engage and buy into the beliefs of AI safety. Public writings, auditable spreadsheets. Sticking with transparency would demonstrate that the movement has integrity and self-consistency.

501c3 public charities are obliged to some transparency already, around financing and executive salaries. And, ~all AI safety orgs abide by the letter of the law. But perhaps, you should be exemplary with regards to its spirit.

It is a duty of powerful actors to be transparent, and AI safety is becoming powerful. Society already asks for transparency from our political leaders, our government, our labs and megacorps. AI safety wishes to influence major actors and pass sweeping regulation. “Dress for the job you want”; prove yourself worthy of this role.

Transparency helps with internal coordination. It scales well. AI safety is about to undergo hypergrowth. It’s no longer 200 people in the Bay Area who all go to each other’s parties. Knowing what other people think and are doing, regardless of who is in or out of particular group chats, will be key to scaling up.

Transparency helps with external recruitment and fundraising. To date, a lot of AI safety people have come to the movement via public writings and transparent reasoning. Many more great people will join as well, if you continue this way. (If transparency is good for recruiting, does that mean that only public comms-focused orgs like 80k and Bluedot should be transparent, versus our research and policy orgs? I’d argue no.)

The rise and fall of “Open Philanthropy”

Any accounting of transparency in AI safety must begin with the saga of Open Philanthropy.

Givewell began as a strong force for transparency among nonprofits. Holden and Elie published constant updates, maintained a listing of their own mistakes, made their spreadsheets available for public critique, and engaged in good faith with commenters. Take a look at the Givewell Blog circa 2007 for a taste of this; I find their attitude beautiful and inspirational. This transparency was key to the early EA movement, and I believe this helped to convince Dustin Moskovitz and Cari Tuna to start funding Givewell with major amounts of money. Together, they started Givewell Labs to explore causes beyond global health, which became the behemoth Open Philanthropy.

Sadly, this golden era did not last. In 2016, Holden published a major update on how they’re thinking about openness (mostly: less open, due to its costs). Beyond their stated reasons, I suspect that once Good Ventures (Dustin & Cari) became more committed to funding OpenPhil, OpenPhil just had less pressure to continue making its thinking visible to the broader public.

From there, things only became less transparent. Fewer public spreadsheets and debating in comment sections. I expect Holden remained a major driver for transparency; I greatly appreciate Cold Takes for this. But Holden left OpenPhil in 2024. And then in 2025, OpenPhil gave up on being “Open Philanthropy” altogether and renamed itself to Coefficient Giving.

After some reflection, I consider this sequence of events to be a failure of principled thinking. I hesitate to say this because Holden himself was (and remains) one of my personal heroes; he’s certainly in the top 5 of people who have shaped my views. (Others include Scott Alexander, Paul Graham, and Jesus). But: I think that Holden and others didn’t properly honor the role of transparency, in the growth of Givewell and later OpenPhil. Phrased extremely uncharitably: being transparent is what got CG to where it is; to discontinue it now is a bait-and-switch.

Does CG itself owe transparency to anyone beyond its funders? I tend to think yes. It owes it to the AI safety ecosystem: CG also draws from a scarce talent pool for its hires, and solicits lengthy application processes from its grantees. And obviously, CG operates as a public charity aiming to benefit the world. If you aim to serve the public, you should engage in dialogue with it.

(One could make a libertarian-ish argument that CG is free to only publish whatever it pleases. I’m sympathetic to this kind of reasoning; in that case, I wish that donors, talent, and grantees would vote with their feet. Hence this essay.)

Beyond internal transparency, I think CG should push the movement for transparency given its outsized role in the ecosystem. CG has been the biggest funder of AI safety, representing more than half of all dollars moved to date, and is also poised to grow rapidly given Dustin’s investments and others (eg Anthropic employees). And, while CG disclaims leadership of either AI safety or EA, I see this as a dereliction of duty; its funding has made these movements what they are. Grantees, and the broader movement, take cues from how CG itself behaves.

Other times when I’ve been disappointed by lack of transparency

Times where I have appreciated transparency

This is a shorter list than the above, but to be clear, I think AI safety does much better on transparency than almost any other comparable movement or community. I critique the movement because I have hope that they will improve.

Manifund’s stance on transparency

To start, the basic premise of Manifund is transparent & fast grant applications. Anyone can post an application on the public internet; fund an application they like; comment about the merits or demerits of a particular proposal; see where and how money moves.

To practice transparency ourselves, we make our source code open; our data wholly available; our finances for anyone to inspect. Beyond our website and newsletter, an unusual amount of our thinking is available on our public Notion.

We’ve recently built tools to increase the transparency in the AI safety & broader EA movement. Trace is our attempt to improve funding transparency by listing every single grant made in the space. The AI safety funder bulletin is our attempt to explain what different funding orgs are up to.

We hope to go farther. We’ll soon be publishing the salaries and roles of all Manifund & Mox staff. We hope to publish more of our own thinking: on funding, AI safety, and what a flourishing future might look like. Movement-wide: we hope to publish analysis of specific orgs, funders, or grants, similar to Givewell of old.

Reasonable steps for transparency AI safety people should consider now

I think all of these are pretty unobjectionable:

Extreme radical transparency (which, might still be good)

 

An ideal version of Manifund might do these things; we’re not there yet, unfortunately.

(To be clear, I’m aware this takes you pretty far down the train to crazytown or a dystopic surveillance state, and there’s probably only like 10 people in the world currently who think this would be a good idea. Consider these points as thought experiments about what might be possible, as opposed to a recommendation to actually implement any of these.)

Objections to transparency

1. It empowers your enemies

True. But it empowers your friends as well. Truth is an asymmetric weapon; if what you are working on is good, you will have more friends than enemies. (If what you’re working on is not good, you should hope to learn about this fact.)

One microscopic example of this: openbook.fyi (a 2023 side project of my now-wife Rachel Weinberg) led EA critic Emile Torres notice and tweet about funding patterns, helping Eliezer update against Slime Mold Time Mold.

2. Sometimes it leads to worse outcomes

True. But I think a general policy of being transparent always will lead to better global outcomes, for AI safety (and, for the world).

Also, I think the argument for transparency comes from more virtue-ethics than consequentialist grounds. And there’s a consequentialist-y reason to engage with virtue ethics: virtue ethics works!

3. Infohazards exist

Okay, I might be willing to hear this out for things like “how to engineer novel pathogens”. But I think this is sometimes used as a justification for being, idk, lazy about what you publish and write.

4. Reputational hazards exist

(I’ll leave this one as an exercise for the reader.)

5. Transparency is costly

True. And I understand that this is a big part of why OpenPhil/CG dialed down its transparency. I just think that the benefit outweighs the costs.

Also transparency doesn’t have to be quite as costly, if you bake it into your operating principles from the start. It feels self-aggrandizing to harp on this, but: Manifund has been open source from day 1, with its proposals and grants always open to the public. We write about our actions, engage with public criticism, and aim to reason in public and update where we’re wrong.

6. AI safety is already pretty transparent

True! And, this is again part of what I loved about the movement. I wish to see this transparency continued and amplified.

Even if AI safety is already doing better on transparency than its reference class (other nonprofits, other researchers, other political actors, labs), I think AI safety must take it upon itself to be a paragon of virtue, given its ambitions.


New Guy @ 2026-09-22T08:21 (+2)

The objections section feels pretty strawmanny here.

1. It empowers your enemies

True. But it empowers your friends as well. Truth is an asymmetric weapon; if what you are working on is good, you will have more friends than enemies.

In the linked essay, truth is called "assymetric weapon" in the context of a logical debate. Unfortunately, the vast majority of the world doesn't form opinions based on logical debate. Storytelling and rhetoric prevail, especially in the short term.

For storytelling, truth (or rather data) is also asymmetric but in the exact opposite way. The more you share (especially of "raw" internal comms), the easier it is to cherrypick parts that make you look bad.

Debunking misleading claims takes significantly more effort than making those claims in the first place. Just take a look at the rise of alt-right or anti-vaccine movement.

 

The objections you stated as 2, 4 and 5 are logically flowing from this:

(Objection 1) Your enemies can use your transparency against you  ->
(4) this can damage the reputation of the org and anyone who works with it  -> a mix of two outcomes happens:

I think that especially the last point is severely underestimated by calls for radical transparency. To be efficient, internal comms hugely rely on shared context, mutual good faith and short loop for asking clarifying questions. External comms have none of those benefits so crafting them takes significantly more effort - one needs to explain the context, pre-empt questions and account for the possible bad faith interpretations.

huw @ 2026-09-22T06:29 (+2)

I am very big on transparency, I strongly agree with this stance! (Admittedly it can be hard on a practical level sometimes). I am thinking more about how charities can open their data up. For example, at Kaya Guides we have our live metrics posted to our website, and I have had a lot of people I wouldn’t normally hear from reach out to tell us how much they liked it. So I think there are also some positive gains to be made there!

I would also note that another objection to transparency is regulatory hazards. This is quite common in the global health field. Sometimes it can be hard to justify transparency today in case a future risk becomes apparent and you might retroactively regret transparency.

Austin @ 2026-09-22T07:43 (+2)

Thanks! Cool to see Kaya's data; I've always been proud of pushing Manifold to make its stats page public as well

Can you say more about regulatory hazards in global health, and/or give an example? I'm imagining something like "you publish what you did in a developing country, then the regime changes, and then they notice you did a bunch of things they don't like, and put you in jail"