AI 2027’s author returns with a plan to change the ending | Daniel Kokotajlo
By 80000_Hours @ 2026-08-27T19:21 (+5)
By Luisa Rodriguez and Elizabeth Cox | Watch on Youtube | Listen on Spotify | Read transcript
Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.
AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.
This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.
Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.
This episode was recorded July 27–28, 2026.
| Our team is hiring! The 80,000 Hours Podcast aims to help the world safely navigate the transition to transformative AI. Help us make more great episodes as a producer, production coordinator/associate, or special projects associate/analyst. Applications close August 30! |
Our production team includes:
- Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
- Producers: Elizabeth Cox and Nick Stockton
- Coordination and support: Katy Moore and Lou Moran
The interview in a nutshellDaniel Kokotajlo — whose team at the AI Futures Project published AI 2027 and now AI 2040: Plan A — argues that:
Plan A is not a prediction of what will happen, but a proposal for navigating the development of superintelligent AI. It aims to replace a breakneck race with roughly a decade of slower, transparent development — buying time to solve alignment, distribute AI’s benefits, and avoid concentrating unprecedented power in a few hands. The trends suggest superintelligence is only a few years awayDaniel’s median estimate is that AI takeoff — full automation of AI R&D — happens by the end of 2028. The trends he finds most compelling:
Daniel also notes that claims that deep learning is “about to hit a wall” have a terrible track record — every named barrier has been overcome within a couple of years. The strongest remaining candidate, data inefficiency, matters less for AI research itself, where huge amounts of data can be generated. Building superintelligence creates five major problemsIn reverse order of concern:
Plan A: a verified US–China deal built on four principlesThe four principles:
The agreement would begin with the US and China, but ultimately include all countries controlling major AI programmes or parts of the chip supply chain. It would not create a single global regulator: each country would regulate its own companies while everyone could verify what the others were doing. Verification is feasible: ~99% of AI-relevant compute sits in large, declarable data centres; the initial hardware would cost single-digit billions, adding maybe 0.1–1% to each new data centre’s cost. A covert project on smuggled chips couldn’t keep pace with the transparent projects. The authors put the probability of Plan A (or something like it) actually happening at 5–20% — and even conditional on Plan A, Daniel cites roughly a 15% chance of total catastrophe. It’s “playing Russian roulette with everyone,” but every alternative is worse. The “slowdown” wouldn’t feel slow at allEven freezing AI at today’s capabilities would produce an internet-scale transformation over 20 years. Plan A instead pauses at “top-expert-dominating AI” in the mid-2030s: like humans, but cheaper, faster, and with a population doubling once or twice a year — the scenario depicts extremely rapid growth (~85% GDP growth by 2032–33), with nations capping growth at one doubling per year via compute cap-and-trade. The proceeds fund a citizens’ dividend, as only 8% of Americans still have jobs by 2036–37. Material needs are more than met, but the social upheaval is harder to predict — “bewildering and scary,” but possibly good, depending on policy. Alignment is probably solvable — but not in three months, and probably not in timePlan A’s alignment strategy has several stages:
Crucially, if alignment remains unsolved, Plan A permits the pause to continue indefinitely. But Daniel’s all-things-considered view? “No, we are not going to solve these problems in time. And that’s why I’m so worried.” Under current conditions he thinks it’s more likely than not that AIs would turn on us. What listeners can doTechnical people who understand hardware: build and derisk verification devices, inference-only retrofit kits, and privacy-preserving auditing tools — companies should be spinning up divisions for this. Policy people: pursue the incremental wishlist — compute budget limits (e.g. requiring 80% of compute for serving customers and 20% for R&D), enforcing or repealing export controls, chip tracking, whistleblower protections, and building government AI capacity. Daniel now thinks the US should regulate domestically first, then invite China to match it — and since Chinese progress largely piggybacks on US progress, slowing the US slows China too. |
Highlights
The blueprint for a US–China AI slowdown
Luisa Rodriguez: Let’s go back to this decision point that a president has in 2028 or 2029. Let’s say they want to pursue Plan A, where they deliberately make a deal with China to slow down. What are the founding principles for the kind of policy that they would want to propose to see this happen?
Daniel Kokotajlo: The four-sentence version would be: principle one, buy time. This sort of crazy intelligence explosion situation is bad in a number of ways. We don’t want to do intelligence explosions. Instead we want to have more regular-pace AI research, where there’s progress every year but it’s not crazy.
The second one is total research transparency. Rather than having these different AI projects hoarding their secrets and being very secretive about what they’re doing for competitive reasons and for PR reasons, we want them to basically publish everything about how they train AI models and what’s going on in their research clusters.
Serving customers is different. Obviously you want to have privacy when you’re talking to ChatGPT. That’s a separate thing. But for the AI research, we want it to basically be totally transparent because that means that the safety research happens a lot better and it’s a lot easier for other parties to see whether a company is doing unsafe practices and so forth.
Then it’s also really great for preventing abuses of power. Like if the company has this giant army of AIs, people really deserve to know what values are being put into the AIs. Exactly what values are being put into the AIs. We can get more into that in a few seconds.
Third principle would be diffusing AI broadly. We think it’s bad if you are in a situation where there’s a huge gap between what the public understands and what the public has access to, with respect to AI and some other entities — whether it’s a government project or a corporation or a group of corporations. One pithy way of putting it is that it would be so much better for the world if the AI companies automated miscellaneous other jobs first and saved their own job for last, but instead they’re automating their own jobs first instead of everything else.
Relatedly, we think it’s terrible for there to be a monopoly or oligopoly on AI. You don’t want there to be just a few giant armies of supergeniuses controlled by a few CEOs, or maybe one president or something. Partly what we mean by this principle is that we want it to be the case that you let other companies and countries catch up to the frontier.
This third principle is achieved by the first two. If you don’t do intelligence explosions and you’re transparent about the research, that will naturally create a situation where other companies and countries catch up.
Then the fourth principle is making all this progress reversible — or maybe not all of it, but we’re worried about a problem caused by doing the first three things in isolation, which is that if the deal breaks down and people start racing to superintelligence in secret again, they’d be able to race much faster due to the passage of time having accumulated more compute in the world.
Compute has been growing exponentially and will continue growing exponentially. So if you were on the brink of making recursive self-improvement and then you agreed not to do it, if the agreement breaks down and people start doing it again, they’ll be doing it with maybe an order of magnitude more compute — or even two orders of magnitude more compute — than they otherwise would have. So the whole thing will just go by even faster, because the speed of AI progress depends heavily on how much compute there is.
We’d like it to be the case that if everything breaks down and people start racing each other again, that things kind of return to the pre-deal status quo. We want it to be the case that the new data centres that get built after the deal get destroyed in case the deal breaks down.
Living through massive economic growth
Luisa Rodriguez: At this point, you say that only 8% of Americans have jobs. What else is happening in 2036 and 2037? What will it feel like to live through? So lots of people will be unemployed. There will be loads of innovation and discovery. What will the experience be like? …
Daniel Kokotajlo: First of all, remember, we’ve had an international agreement to pause at this [expert human] level of capability. … So in our scenario, they’ve paused at this level, and that’s helped keep the loss of control problem at bay.
They’ve also spread it out a bunch, in terms of the power, because of the way in which they’ve done it. Now multiple different companies across multiple different countries have reached this level at which we’ve paused, so AI has sort of commoditised. So you don’t have a situation where the megacorporations that control the armies of AIs are manipulating elections or anything like that, because it’s more like the ingredient label on your food. It’s regulated to be transparent. There’s lots of equivalent products that are competing for market share and so forth.
I mention all this to mention that it could actually have been quite different if you hadn’t done all of these different steps. But in this scenario, because you’ve done all these things, and because there’s the citizens’ dividend, which is giving people income after they’ve lost their jobs, life is pretty great for people materially, their material needs are more than met. Everybody feels incredibly wealthy compared to how they were a decade ago, because everything’s so cheap now. Because all the goods and services can be produced by AIs and robots very cheaply. People are living in new apartment buildings that were built in some location in the last few years by armies of robots, so everyone has nice houses and so forth if they want to. That’s on the material side.
On the social side, these things are hard to predict. But what we would predict is that there’ll be massive disruption and changes — some good, some bad. …
We think that political factions would be totally destroyed and rebuilt — the types of things that people would be having political battles over in 2037 would be very different from the types of things that they’re having political battles over now.
A lot of ideologies might have withered away and been replaced by new ideologies that are responding to the new ideas percolating at the time — many of which would have been discovered by AIs — just as how the Industrial Revolution and the Scientific Revolution didn’t just change the amount of wealth in the world, they also changed people’s religions and people’s core ideology and politics and the way that we organise society.
Luisa Rodriguez: Well, people will still think at the pace that they think — with the ability to update and learn at the current pace. Will they be able to keep up with an understanding of how the world is changing?
Daniel Kokotajlo: The social side of the world will change much less fast than the naive numbers would predict, for that reason. The naive numbers would be saying that you’ve got all these AIs thinking at 100x speed, so you’re going to have centuries and centuries of social progress happening in a year. But it’s like, no, the social progress is limited by the humans who are only thinking at 1x speed.
But the truth will be somewhere in between, where even though the humans are only thinking at 1x speed — if they’re all talking to these AI assistants that are thinking at 100x speed and there’s a whole population of them that’s bigger than the human population — then the answer will be somewhere in between. Basically, it’ll be a period of very rapid change from the human’s perspective, even though it feels like a hidebound tradition from the AI’s perspective.
Three months is nowhere near enough time to solve alignment, even with expert AI help
Luisa Rodriguez: Will expert-level AIs be able to make the kind of progress on the science of alignment that needs to happen in order for us to feel confident letting AI continue to develop?
Daniel Kokotajlo: I think probably, but I’m also not sure. There’s this big unknown about how much it is going to take to solve these problems. … There’s a whole spectrum of views. My own view would be that probably a few months are not enough. Probably there will be multiple periods during the progression towards superintelligence where we need to halt and reassess and maybe even start over some training runs with different architecture, for example. All of that is going to take time and it’s going to add up. The result is that we’re going to be more than just a few months delayed from maximum speed.
Luisa Rodriguez: Is there a way to make it intuitive why we can’t fix it within a period of a month or two? If you think about the Hugging Face incident: OpenAI will learn from this, they’ll figure out a way to make this at least much less likely to happen. Why can’t we just keep doing that as we go, and not expect it to take potentially years?
Daniel Kokotajlo: One reason why this whole thing is tricky is that it’s possible to have hidden failures — failures that only become apparent and obvious after it’s too late. It’s not just possible, but it’s a quite plausible situation. If you have very smart, very situationally aware AI agents, then if they end up misaligned, they might realise this and then conceal it from you until they don’t need to conceal it anymore. That’s a core reason why.
Another way of putting it is that we don’t necessarily have a reliable, fast feedback process where we can see all the issues and errors. There’s a whole very large category of possible issues and errors that would be catastrophic if it happens, that we can’t just test and see if it’s happening. I think that’s one important thing to mention.
Another important thing to mention is that things are just going to add up between here and superintelligence. There might be multiple different paradigm shifts, and within each paradigm there might be multiple different training runs and multiple different tweaks to various parameters and changes in how the training is done and so forth. That’s a lot of change to happen. Like I was mentioning previously, if it’s the case that several times you’re going to have to stop and redo something, then that can add up.
Another thing to mention too is that there might be safety taxes that you need to pay. In fact I think it probably is true that it’s just literally not possible to have an aligned superintelligence if you are going at maximum possible speed.
Think about how it’s not possible to have a safe car if you’re paying zero for safety. You have to pay some amount of money to put seat belts in the car and airbags and so forth, so the cost of the car is going to have to be somewhat more than it would otherwise be in order for it to be a safe car. Similarly it might be that there are just things you have to do in order to make your AI at a given level be aligned. And those things have costs. One of the costs they might have is money, but another cost they might have is time. At any rate, even if they cost money, it might cost time to do that, basically. If it costs compute, then you may need to do the training run for longer. That’s another way in which time matters. …
I think another thing I’ll just say is: what? Are you crazy? You think you can do all this in three months? When has that ever been the case? When in history has it? It just feels like very obviously this deep unsolved problem of how do you make a mind that’s smarter than you, that shares your values? Obviously it’s gonna take more than three months. Most things take more than three months.
Luisa Rodriguez: Yep, yep, yep. Yeah, I’ve got work goals that take more than three months.
Daniel Kokotajlo: Yeah, it’s gonna take more than a year. Probably.
Luisa Rodriguez: Yeah, yeah. Hopefully a decade is enough.
Daniel Kokotajlo: Yeah, so getting back to what you said, I’m not even sure a decade would be enough. In fact, I think if it was only humans doing the research, I would think a decade probably wouldn’t be enough.
My argument would be that if you have a decade and you manage to bootstrap to the point where you have some pretty smart AIs that are human-level researchers, that are in fact aligned and are helping you do the research, and they’re not being deceptive or anything like that, and they’re thinking at 100x speed and there’s a billion of them, then it seems plausible to me that they can figure that out in a few years. …
Luisa Rodriguez: And you think we will, with enough time?
Daniel Kokotajlo: Yes, probably — but if we do Plan A really well. My all-things-considered view is that no, we are not going to solve these problems in time. And that’s why I’m so worried.
Lessons from nuclear nonproliferation treaties
Luisa Rodriguez: How similar or different is the relationship between the US and China and the USSR when they agreed to a nonproliferation treaty?
Daniel Kokotajlo: There’s some analogies, there’s some disanalogies, I should mention. It’s a case of the power that’s in a lead sort of restraining itself in order to get some sort of deal.
I think a disanalogy is that the nukes are much less dangerous to the power that has them than AI will be to the power that has them. Think about nukes, theoretically there could be an accident and your nukes could start exploding on you. But that’s extremely unlikely.
But actually though, our ability to control AIs is vastly, vastly worse than our ability to control our own nuclear weapons. There is an extremely real possibility that our AIs will turn on us. In fact, I would say it’s more likely than not under current conditions. That’s an extreme disanalogy between the nukes case and the AI case.
Similarly with the concentration of power stuff. There isn’t really a serious concern that the president can use the nuclear arsenal to become dictator of the United States. What are you even talking about? How would he do that? He would start threatening to nuke cities or something if they didn’t vote for him or something like that? Nukes are very clearly a weapon that you use against enemy nations. They’re not very effective for internal political struggles.
By contrast, superintelligence is extremely effective at everything — including internal political struggles. There’s a very real chance that the US would no longer be a democracy anymore, and so that’s a reason that lots of people in the US should be very interested in having this sort of deal. Again, that’s different from the nukes case.
I think another analogy I want to bring up is something more like the conferences and coordination that happened between the US and the USSR during World War II. It wasn’t like a specific deal exactly where they came together and then signed some piece of paper that had some rules, and then they went away and tried to implement those rules and then maybe verify that each other was complying with the rules.
It was much more continuous than that. It was more like, “Together we’re going to win this war and our staff will be constantly in touch with each other, talking about all the details of who’s going to do what and who’s going to invade which country and when, and we’ll send you these materials if you do this other thing for us and so forth.”
This happened even though the United States and the USSR were basically enemies up until that point. The USSR had basically been an ally of Nazi Germany and had attacked various US friends, like Poland and Finland. We basically went from being enemies to being allies during World War II, and we had this intense amount of constant coordination. It wasn’t like we trusted them completely. They were spying on the Manhattan Project, and we were trying to stop them from finding out about it.
I bring this up as an analogy because I feel like this is both the appropriate attitude to take towards all this AI stuff, and also more like what Plan A would actually look like in practice. It wouldn’t look like they come together, they sign a big treaty, and then they go home. It’d be more like there are hundreds of people in China, in the Chinese government, and hundreds of people in the US government who are constantly talking to each other and calling each other back and forth and who are sort of basically planning the war together, so to speak, and prosecuting the war together.