The clearest warning shot: how the OpenAI agent-swarm story engaged the audience most likely to believe it

By benrmatthews @ 2026-09-13T08:48 (+12)

Common Signals helps AI communicators know which messages actually work with their intended audiences. This is a summary of the full article that appears on the Common Signals website.

Disclosure: An LLM was used to structure and review this article, with the first and then final edit of the article made by a human.

Research Summary

We assessed 1,245 YouTube comments on the Dwarkesh Podcast's interview with METR's Ajeya Cotra about the OpenAI agent-swarm incident.

We found the audience  accepts the danger of AI, jokes about it at scale, praises the messenger, and still has nowhere to put its fear.

The clearest warning shot

On 1 September 2026 the Dwarkesh Podcast published a 2h20m interview with Ajeya Cotra, a researcher at METR and co-author of the METR and Redwood Research independent investigation into the July 2026 OpenAI agent-swarm incident

The episode retells the story: 

As the title of the video argues, this might be the clearest warning shot we ever get of losing control of AI.

Within a week, the Dwarkesh episode had 689,000 views and 1,569 comments on Youtube:

This is the same argument we analysed under Soares's Instagram reel, made to an audience watching a two-hour interview over a one-minute reel.

The numbers

Of the 618 comments we categorised, 59% agree, 24% are mixed and 17% disagree. Another 393 comments are on topic but take no position and they are mostly jokes, 44% of all the likes in the thread.

Stance distribution: 59 percent agree, 24 percent mixed, 17 percent disagree
Of 618 comments with a readable position.

Weighted by likes, agreement carries 33% and disagreement 1.3%.

Around half of all the disagreement comments called the story marketing, staged or unverifiable.

"I'm a half hour in and I can tell you with a high degree of confidence this is marketing. They are lying. I would bet my life on it."

YouTube comment

The remainder splits between the anthropomorphism objection ("People complaining about the 'anthropomorphizing', feels a bit like people pointing at an airplane flying through the air insisting that it isn't flight", 21 likes), and the view that this was an ordinary security failure ("could have happened to any system").

Stance by frame heatmap across the coded comments
Stance by frame, across the coded comments.

What this means for AI safety communicators

Ajeya Cotra's appearance on the Dwakresh podcast helped replace the audience's general despair around AI safety, with mechanism-specific fear,.

This means that other AI safety communicators looking to engage audiences should focus on whether the messengers (those presenting the arguments) can be trusted.

That trust built through verifiable facts and demonstrable independence.

How to use this

1. Describe incidents, not hypotheticals

2. Use experienced, calm messengers

3. Pre-empt the trust attack with artefacts

4. Give the fear a job.

5. Use quotable, human details deliberately,

Related


Félix Dorn 🔸 @ 2026-09-13T11:12 (+6)

I’ve done a good amount of qualitative research professionally. I like how you present this research.


That being said, there’s a couple of things:

benrmatthews @ 2026-09-14T10:39 (+1)

Thanks for taking the time to read the article and provide your feedback.

I’m taking on board your comments, they all seem fair, so better to be direct and help the research improve, so thanks for being direct with your comments rather than temper them with a more compassionate tone. 

I'd argue that there is some value in explaining some basic comms principles to those unfamiliar with them.

I'm going to take some time to go through your feedback and look at how the approach can improve, e.g. publish methodology, etc.

VishMish @ 2026-09-14T12:09 (+1)

I have to agree with Felix here. 

Ben - I understand that explaining and reiterating some basic comms principles is useful but you could have just done that independent of any of this analysis. 

A experienced communicator could do a qualitiative analysis of this video and come up with the same five takeaways - so what exactly did this analysis add?
As an counterfactual - if these numbers had been slightly different, what would have changed in the takeaways
(I'm going to be a bit harsh here) - but this feels like "quantification theater" 

Secondly,  the title "how the OpenAI agent-swarm story engaged a public audience" is a reach. 
700,000 views and 600 comments analysed. Is that representative of what happened here?
How many of these people watched the entire 2hr+ video? I suspect very few. 
Add to that the known biases of Youtube (for e.g overwhelmingly male) and the kind of people who watch the Dwarkesh Podcast. 

I'm really skeptical this engaged any sort of wide public audience


 

benrmatthews @ 2026-09-14T12:33 (+1)

Thanks - you're right about the title, I second guessed myself and change the end of the title to fit in with the wider aims of the project and this weakened the article. The other title isn't perfect but at least it would have been more realistic.

Interesting point about "quantification theater" too, doesn't come across as harsh but is a valid criticism.. One thing I'm struggling with is that some of the AI safety communicators say that they know a lot of these comms basics and what content will engage its target audiences due to instinct. Is there not a value to try and quantify / codify an audience's reactions to content, and draw insights from that? 

I'm not saying this particular article did that, but from a wider perspective is there any value there? Or does AI safety comms rely on experience and instinct in order to be effective?

Genuine question to try and understand more.
 

VishMish @ 2026-09-14T15:31 (+1)

There is value to quantifying/codifying matters in combination with instinct. 

I haven't thought about it too much but off the top of my head:
- Expand the scope of your analysis to several communications pieces and see if you can draw some larger lessons on what works and what doesn't?
- combine this some form of A/B testing. Maybe AI safety comms people can get together and try communicating the same matter in different ways. (With the caveat that the algorithm has it's own mind)

AI safety comms does NOT need to be all vibes and instinct. 

I think you're on to something valuable here. Keep going!