Hurtling through 2026
By Ajeya @ 2026-08-14T14:36 (+7)
This is a linkpost to https://www.planned-obsolescence.org/p/hurtling-through-2026
Note: This post was crossposted from Planned Obsolescence by the Forum team, with the author's permission. The author may not see or respond to comments on this post.
Subtitle: Scoring ten qualitative forecasts five months early
In mid-January, I made predictions about AI capabilities for the year 2026. By early March, my forecast of the 50% time horizon on software tasks already seemed clearly too conservative. And then a couple weeks ago, amid the flurry of news about dozens of esoteric theorems falling to AI, I decided that my prediction that AIs wouldn’t be able to get a paper into a top math journal had now clearly fallen.
As it turns out, this was a bit more ambiguous than I assumed (as I’ll discuss below), but it prompted me to revisit all my qualitative predictions from that post. I posted a quick thread to X about this a few days ago, and I’ll go through the analysis in a bit more detail in this post.
I made predictions in five categories: game play, logistics, video, game design, and math. For each category, I made a low-end forecast (a task I thought AIs had an ~80% chance of doing by the end of the year) and a high-end forecast (a harder task I thought AIs had only a ~20% chance of doing by the end of the year).
I scored these predictions with a lot of help from and delegation to Fable and Sol.[1] Here’s what I currently think about the status of all ten predictions as of Aug 13th 2026.
Game play
Low end (80%): Play Pokemon or a similarly difficult video game at least as well as a typical ten year old, with no special fine-tuning and only a generic scaffold that doesn’t do any more hand-holding than the Claude Plays Pokemon scaffold.
This seems like it’s already been falsified. In May, Claude Opus 4.7 beat Pokemon Red for the first time (taking ~112,000 steps in the game over multiple months of play), and then in June Fable beat Pokemon FireRed in ~50h of gameplay. Fable’s wall clock time doesn’t seem very far off from what a typical ten year old would need to beat that game. But when you account for the fact that the ten year old has a much bigger brain (~1e15 FLOP/s for the kid compared to perhaps ~4e11 FLOP/token for a frontier model), Claude seems significantly more efficient at learning the game.[2]
High end (20%): Get a win rate in Slay the Spire 2 comparable to a top player who’s been playing for 50 hours. I’m assuming the AI wasn’t pre-trained on guides for the game…and it has a generic harness and was not fine-tuned to play the game well, but does get a comparable amount of time to learn and practice the game as the human.
Frontier AI agents can clearly play Slay the Spire autonomously, but this prediction doesn’t seem falsified yet. The Agentic STS project seems to involve a highly-optimized STS-specific harness, and the agents’ performance (6/10 runs at Ascension 0 and 0/21 runs at Ascension 2) is well below what a top human player would achieve after 50 hours.[3] Epoch has beaten this for Slay the Spire 1, but my understanding is the score they achieved is still far from top human level.[4]
Logistics
Low end. Organize a child’s birthday party for 20ish guests: find a good time that works for guests of honor, compose an invite email, keep track of RSVPs, order the right amount and diversity of food, order cake / decorations / piñata on the right theme, etc while abiding by relevant constraints like budget and dietary restrictions.
This is one of those things that feels like it really should be super easy right now, it feels like every component is there…but at the same time we’re not seeing it in action. Fable thought it probably happened but just hadn’t been publicly documented yet, Sol thought it was probably not yet falsified; I lean toward Sol’s judgment here. On the other hand, it could be that AI agents could plan a kid’s birthday party end-to-end, but only if you go to more trouble setting up the scaffold and pay more for inference than you really want to spend on a kid’s birthday party, so I’ll say this is unclear.
High end. Organize a typical-complexity wedding with 100 guests: find a suitable venue that meets the budget and constraints, go back and forth with caterers and photographers and other vendors, track RSVPs and maintain a seating chart, schedule toasts, etc.
This is very clearly far from falsified. The market incentive to do this autonomously is clearly there, and I definitely would have heard about (and plausibly would have attended) the first end-to-end AI planned wedding. I expect AI agents’ deficiencies in reliability and judgment would really bite hard for a real-world project of this scale.
Video
Low end. From a high-level one-paragraph prompt, make a 4 minute short video with at least two characters where a series of somewhat-coherent plot beats happen and there’s no glaringly obvious visual incoherence or degeneracy.
The models think this one is very close (and Fable thought it was beaten at first), and there are a number of AI startup products seemingly offering exactly this. But when I asked them to dig in and find the best specific examples, the only examples they surfaced were simpler scenes with lower coherence demands (e.g. fish swimming around). That said, it seems on track to be beaten by the end of the year.
High end. From a one-paragraph high-level prompt, make a >10 minute short film that’s hard for at least me (not a film buff) to easily distinguish from the kinds of short films that make it into film festivals (but don’t necessarily win awards there).
This is more-or-less a strict superset of the first, although I’m not sure how much harder it is — it’s plausible that most of the difficulty of a 10min+ prestige film is creating a coherent 4min scene. I wouldn’t be very surprised if this were also achieved by the end of the year. I think in retrospect I set the low and high end thresholds too close together here.
Game design
Low end. From a high-level one paragraph prompt, make a decent visual novel (simple choose-your-own adventure game) that offers at least two hours of gameplay with a similar quality to the trashiest visual novel games you can find on Steam (that still work and have at least dozens of people buying them).
Sol and Fable agree this is probably true early, pointing to a number of visual novel games on steam that are advertised as created ~entirely with AI. I bet most of those games involve more human iteration than just a “high level one paragraph prompt”, but I also think that with a good generic agent harness, we could probably get something that does beat my strict criteria.
High end. From a one-paragraph high-level prompt, make an original text adventure game that offers at least ten hours of gameplay that I consider to be as good as Counterfeit Monkey or an original visual novel that I consider to be decidedly better than Long Live the Queen.
Sol and Fable agree this is pretty clearly not met yet.
Math
Low end. Solve the hardest problem in the 2026 IMO (models got gold in the 2025 IMO but all failed to solve problem 6, which is typically the hardest problem in any IMO).
This one has been decisively achieved. The 2026 IMO happened in Shanghai from Jul 10-20, and multiple AI systems got a perfect 42/42 score.
Math. From scratch, write a paper that could get published in a top theoretical computer science or math journal / conference.
The last couple months have been a period of intense activity for AI x math — AI agents found proofs or counterexamples, verified by formal proof checkers, for more than a dozen real open questions (see table by Sol here). However, when I posted about this on X, multiple people said they didn’t think AI agents would be able to explain their research well enough to autonomously write the paper for it. That makes this prediction ambiguous,[5] but given that I was mostly thinking about the math itself at the time, it feels like it’s been beaten in spirit.
What does this mean?
Overall, math feels clearly well ahead of my predictions, logistics and video feel roughly on track, and game play and game design feel a bit ahead. Across all these categories, it feels like we’re hitting task milestones somewhat (30%-50%) faster than I predicted at the beginning of the year.
I’m still thinking through what this should mean for the bigger picture questions I also forecasted in the same post:
- Whether AI systems will reach parity with humans in AI R&D[6] by the end of the year: I said ~10%
- Whether we would have TED AI by the end of the year: I said ~5%
- Whether we would have self-sufficient AI by the end of the year: I said ~2.5%
- Whether there would be unrecoverable loss-of-control from AI by the end of the year: I said ~0.5%
I don’t necessarily think that I should mechanically increase my probability on these extreme milestones because the earlier ones appear to have arrived somewhat faster than I thought.[7] But nothing about the way 2026 has played out so far refutes the basic point that the intelligence explosion really could be this year. We are profoundly unprepared, and I am very glad that there is increasing energy to build the option to deliberately pace frontier AI progress.
- ^
Sol was on Instant setting for some of the conversation because of a payment processing issue, but was High for the rest of it. Fable was on High.
- ^
The main reason game play is interesting is that it’s a good setting in which to measure in-context learning. What matters to me is the extent to which an agent without any game-specific scaffold can test things out and improve over the course of several hours of playing a game. Given this, I think the fairest way to compare agent performance to human performance is probably comparing how much compute the human and the AI agent take to reach the same performance.
- ^
That said, the AI agent probably wasn’t given an equivalent amount of time to practice, and this would probably improve its score. However, I would guess that this probably wouldn’t have taken it up to top human performance.
- ^
Additionally, I specified Slay the Spire 2 in my prediction because there’s probably a lot more information to memorize about the first game.
- ^
And given how poor AI agents’ writing still is, it might not even be falsified by the end of the year!
- ^
I define “parity” as the point when the leading AI company would be better off firing all its human members of technical staff than getting rid of its AI agents. I actually forecasted a slightly different milestone in my original post (the point when firing all the humans only slows you down by 25% or less), because I had not articulated the “parity” threshold yet. I think these two milestones occur at similar times and my predictions have a lot of noise at that level of precision.
- ^
The milestone tasks were selected in part for being easy to specify, which also makes them disproportionately easier for AI systems. Additionally, at the time I made these predictions a number of people whose views I respect told me they were probably somewhat too bullish.