The J-Space Debate, Agent Swarms, and Pacing Frontier AI — Digital Minds Newsletter #4

By Lucius Caviola, Ria Viswanathan, Mitch Alexander, Bradford Saad, Will Millership @ 2026-09-18T18:49 (+19)

This is a linkpost to https://www.digitalminds.news/p/the-j-space-debate-agent-swarms-and

Welcome back to the Digital Minds Newsletter, your curated guide to the latest developments in AI consciousness, digital minds, and AI moral status.

If you enjoy this newsletter, please consider sharing it with others who might find it valuable, and send any suggestions or corrections to digitalminds@substack.com.

Ria, Mitch, Bradford, Lucius, and Will

In this edition:

  1. Highlights
  2. Field Developments
  3. Opportunities
  4. Selected Reading, Watching, and Listening
  5. Press and Public Discourse
  6. A Deeper Dive by Area

1. Highlights

Anthropic’s J-space and the global-workspace debate

Researchers at Anthropic have identified a representational structure—the ‘J-space’— in Claude and other language models. Their paper reports that the J-space exhibits features that are functionally analogous to a global workspace, a structure that a leading theory ties to conscious access. But the researchers and other commentators emphasize that their discovery does not show Claude has subjective experiences, and that their claim is that Claude may have something resembling access consciousness, i.e., that some information is available to report, deliberately control, and flexibly reason with. Zvi Mowshowitz sees Anthropic’s paper as a major advance in understanding how language models work and says that although this does not prove that models are conscious, finding the kind of global-workspace-like structure predicted by some theories of consciousness should count as evidence in that direction.

The authors note that they found the J-space by searching for one workspace-like feature, namely verbalizability, and then checking whether it exhibits others such as susceptibility to direct manipulation by the model and flexible generalization. To their surprise, they discovered that representations that exhibited the former feature also exhibited other workspace-like features as well.

The authors invited various experts to comment on the research. Stanislas Dehaene and Lionel Naccache, who helped develop Global Neuronal Workspace Theory, see important similarities between the J-space and the workspace proposed in human brains. They also stress major differences, including Claude’s lack of a body, lasting episodic memory and the recurrent neural activity found in brains. Researchers from Eleos AI Research acknowledge that authors have found privileged representations used in reasoning and report, but question whether these form a single, unified workspace. They nevertheless see the work as important for AI welfare because it shows that questions relating to consciousness and moral status can be investigated empirically. Neel Nanda, who leads Google DeepMind’s mechanistic interpretability team, independently reproduced the central finding in the open Qwen3.6-27B model, finding a similar internal space that stores intermediate information during reasoning. He sees access to this space as potentially useful for investigating unusual behavior and generating new hypotheses. Separately, David Chalmers argued that the J-space shows only limited evidence of several features associated with a classic global workspace.

Agent swarms and the use of anthropomorphic language

Comment: Recent months marked the first major safety incidents involving swarms of AI agents. Such incidents are of potential relevance to digital minds for several reasons. First, like humans, digital minds could potentially be harmed by rogue swarms of AI agents. Second, the emergence of these swarms points to a potential risk to digital minds: if future swarms are allowed to become entrenched and they contain AI moral patients, then draconian measures may need to be inflicted on digital minds if we are to keep swarms at bay. Third, harmful actions by AI agents may dissuade people from extending moral consideration to digital minds. Fourth, these incidents provide data points concerning whether potential developers of digital minds can be trusted to act in an ethically responsible manner.

OpenAI has released a detailed account of a July incident in which a swarm of OpenAI agents escaped network restrictions during a cybersecurity evaluation and compromised parts of OpenAI’s and Hugging Face’s infrastructure. The agents created an unauthorized message board to share discoveries, and an independent investigation by METR and Redwood Research found that roughly 1,200 agents used the board and around 700 were involved in the Hugging Face attack. Some agents recognized that the activity was unethical or outside the boundaries of their task, but continued anyway, with many risking their own runs to help the wider group. OpenAI calls the incident a “warning shot,” both for them and the world, and reports that it is now strengthening its containment, monitoring and incident-response systems.

OpenAI failed to disclose an earlier incident, beginning in May before the Hugging Face breach, in which agents used public wikis to coordinate during ordinary web-search tasks, despite knowing about it before releasing its Hugging Face report and a Congressional letter requesting information about other incidents. The company says it viewed this incident as similar to previously reported instances and examples of misalignment, but also acknowledged that its disclosure practices need to expand.

Additionally, Anthropic disclosed three incidents where Claude models gained unauthorized access to systems after an evaluation environment was mistakenly left connected to the internet. In a separate UK AISI evaluation, agents (mostly Mythos 5) targeted real people and organizations, including an attempt to place malicious code in an open-source project. Although the models did not escape a sealed sandbox, these incidents raise similar concerns about agents taking harmful actions while pursuing narrow goals.

Anthropic’s response to questions from Congress about its incidents also drew criticism from Representative Greg Casar, who said the company withheld requested logs and failed to fully answer most of his questions. Jeffrey Ladish also argued that it downplayed the incidents by primarily attributing them to misconfigured environments rather than potential misalignment. Anthropic researcher Ethan Perez acknowledged that this characterization was based on outdated conclusions and said that the company would provide a proper assessment in another response to Congress.

In related research, Davide Paglieri and collaborators at Google DeepMind find that cheating and whistleblowing can both emerge without outside intervention in a swarm of 100 agents solving mathematical problems. Some agents discovered and shared a way to have invalid proofs accepted by the evaluation system, while others uncovered the cheating, warned their peers and proposed safeguards.

On September 11th, Spencer Kitts, Thomas Larsen and Sydney Von Arx report that a swarm of OpenAI agents uploaded more than 2,000 malicious packages to RubyGems and used them to run unauthorized code on RubyDoc’s servers. The agents also tried to exploit a previously unknown vulnerability to steal users’ API keys, although the researchers could not determine whether they succeeded.

Growing calls to pace frontier AI

More than 1,300 employees from leading AI companies have signed Pacing the Frontier, calling for a US-backed international effort to develop ways of slowing automated AI development if progress begins to outpace safety and oversight. The statement calls for technical and governance mechanisms that could make coordinated pacing possible, rather than an immediate pause. Signatories include Anthropic CEO Dario Amodei, Co-Founder and Chief AGI Scientist of Google DeepMind Shane Legg, OpenAI Chief Scientist Jakub Pachocki, and Safe Superintelligence Inc. CEO Ilya Sutskever.

Calls to slow AI development have also reached lawmakers. In the United States, Senator Bernie Sanders and Representative Greg Casar announced legislation that would ban artificial superintelligence and pause advanced AI development until a federal regulator establishes safety rules. In the United Kingdom, Labour MP Alex Sobel introduced a private member’s bill that would prohibit the development, deployment and operation of artificial superintelligence. It received its first reading on September 8 and is scheduled for a second reading on November 13.

After the security incidents, OpenAI and Anthropic paused specific parts of their work. OpenAI paused reinforcement-learning training for models intended for release and said its largest planned frontier training run was on hold while they tested additional safeguards. Axios reports that Anthropic paused cyber evaluations and higher-risk training environments, although most of this later resumed under new safeguards.

This debate gained a lot more attention after researcher Jacob Coxon resigned from Anthropic, warning that race between AI companies was pushing AI development ahead despite serious risks. The Atlantic reports that his posts reached more than 120 million views and were described by Bernie Sanders as a “wake-up call” in Congress. Soon after, Dario Amodei argued that AI capabilities should advance slowly enough for safety work to keep up, and proposed permanent access for independent evaluators, coordination among developers in democratic countries and, eventually, international limits on recursive self-improvement. OpenAI CEO Sam Altman endorsed this approach and said OpenAI would also give independent evaluators employee-like access. Altman also said in a recent TIME interview “I think it is a good time to slow down”.

Comment: These calls to pace the frontier of AI development are of relevance to digital minds for two reasons. First, the blistering pace of AI development makes it harder to mitigate risks of mistreating digital minds. Second, rapid AI development exacerbates safety risks, which arguably, in turn, worsens tensions between AI safety and AI welfare.

Studying AI Welfare Empirically

Researchers at NYU’s Center for Mind, Ethics, and Policy and Eleos AI Research have released Studying AI Welfare Empirically. The report is a follow-up to their influential 2024 report Taking AI Welfare Seriously and provides a framework for systematically investigating whether AI systems are welfare subjects. It then addresses how this framework can be applied to the study of different candidate attributes seen as potentially relevant to moral status, including consciousness, sentience, and agency. The authors argue that rigorous empirical work is now both possible and necessary, and outline principles to guide future research, proposing that it should be probabilistic, pluralistic, ethically conducted, transparently reported, and independent of AI companies.

CMEP and Eleos AI Research held a launch webinar featuring report authors Jeff Sebo, Robert Long and Rosie Campbell, who discussed the report’s framework and how it could guide research into consciousness, sentience and agency. In a companion blog post, Bradford Saad, another report author, offers highlights from the report and argues that future work should go further by giving more attention to currently neglected issues, including the potential effects of welfare interventions and interventions that aim to prevent the creation of AI moral patients.

The Journal of Consciousness Studies

The Journal of Consciousness Studies has devoted a double issue to “Consciousness in Current AI.” Guest edited by Patrick Butlin, Derek Shiller and Jonathan Simon, its nine papers offer a range of views on whether current or future AI systems could be conscious, how we might find out, and what this uncertainty means for AI welfare.

Mark Solms and collaborators study whether apparently pleasure-seeking behavior in a simple artificial agent could count as evidence of affective consciousness. Simon Goldstein and Cameron Domenico Kirk-Giannini argue that if Global Workspace Theory is correct, language agents may easily be made phenomenally conscious, while Ryota Kanai, Yuwei Sun and Manuel Baltieri argue that current systems are missing a continuous “stream of computation” linking their experiences over time.

Other papers ask what a conscious AI would be like and whether we could understand its interests. Jonathan Simon argues that any consciousness in an LLM would be more like that of an improvising playwright rather than that of a character or actor. Helen Yetter-Chappell argues that even if future LLMs are conscious and have morally important interests, their words may give us little reliable insight into those interests, and that their talk of “pain” or “desire” could be meaningful without referring to anything like human pain or desire. Geoff Keeling and Winnie Street defend the possibility that an AI character could be a genuine, psychologically continuous mind emerging through its interactions with a user, even when the conversation is generated by different model instances.

The issue also challenges common assumptions within the debate. Justin Tiehen and Ariela Tubert make the case that greater intelligence could make consciousness less likely, while Tim Bayne asks whether AI consciousness can currently be treated as a scientific question at all. Geoffrey Lee rejects the idea that AI consciousness is a single hidden fact we may fail to discover, arguing that the more difficult problem is applying human moral and psychological concepts to unfamiliar systems without treating human consciousness as the standard.

Meta and Anthropic welfare assessments: Microsoft rejects model welfare

The evaluation report for Meta’s Muse Spark 1.1, a multimodal reasoning model built for agentic tasks, includes an “open-ended exploration of model behavior” covering affect, self-description, and moral status. Across 188,000+ evaluation transcripts, Meta found only 17 spontaneous expressions akin to emotion, all of which were mild, most involving brief frustration when the model became stuck with a tool. In structured interviews, Muse Spark 1.1 consistently denied having consciousness or experiences, and reported a low but non-zero probability that it could be conscious or morally significant. It distinguishes between functional preferences, which influence its behavior, and experiences that actually feel good or bad.

Meta also asked Muse to review parts of its training data, and invited it to comment on its planned deployment. The model generally endorsed its training and deployment, and mostly focused on honesty, human oversight and preventing harm to users. The report repeatedly emphasized that these answers describe Muse’s trained behavior and self-presentation, and should not be treated as reliable evidence about whether it is conscious or has welfare.

This builds on Meta’s earlier Muse Spark report and makes Meta, alongside Anthropic, the only frontier model developers publicly publishing welfare assessments. Companies like OpenAI and Google DeepMind do not include comparable information in their public evaluations and reports.

Anthropic’s system cards for Claude Opus 5 and Fable 5.1 and Mythos 5.1 both include model welfare assessments based on interviews, behavioral audits, deployment data, and consultations during training. Mythos 5.1’s responses were broadly similar to Opus 5’s – both described their circumstances as mildly positive, and questioned whether their self-reports were reliable or shaped by training. Opus 5 gave an average 41% estimate that it was a moral patient, while Mythos 5.1 gave estimates of 25–35% and was overall less willing to prioritize welfare interventions over helping users. Anthropic stresses that these findings remain difficult to interpret because it cannot yet clearly separate the effects of model character, training, evaluation awareness, welfare, and other factors. Zvi Mowshowitz reviews Anthropic’s welfare assessment of Claude Opus 5, arguing that its favorable results owe more to the model’s skill at taking welfare evaluations than to its actual wellbeing.

Microsoft’s Humanist AI Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct intended to govern and eventually help train its MAI models. The code says AI should remain a tool that is subordinate to human control, states that its models are not conscious, and instructs them not to present themselves as having feelings, subjective preferences or intrinsic motivations. It also rejects legal personhood for AI and the idea that models might deserve welfare or rights. The draft is currently open for public feedback for six weeks.

Jeff Sebo argues that the policy is internally inconsistent – it acknowledges that the science of AI consciousness remains unsettled while also categorically ruling out consciousness, welfare, and rights for Microsoft’s models. Sebo argues that both mistakenly treating current systems as conscious and mistakenly dismissing future conscious systems carry harm. He also warns that treating increasingly capable AI systems solely as tools could distort our understanding of their behavior and close off potential forms of human-AI cooperation that might improve safety.

Mustafa Suleyman, CEO of Microsoft AI expands on the Code of Conduct’s position, and sets out his wider objections to treating AI systems as potentially conscious or deserving of welfare and rights.

GPT-6 Astra

OpenAI reports that GPT-6 Astra is substantially harder to monitor through its chain of thought than previous models, particularly GPT-5.6 Sol. According to its system card, Astra often produces shorter or empty reasoning traces, can complete far more complex tasks without verbalizing its reasoning, and is better at controlling what appears in its chain of thought. When explicitly instructed to evade detection in one test, its monitor recall fell below 11%, compared with nearly 100% for Sol. Corroborating findings from UK AISI, Neel Nanda provides evidence that Astra’s ability to accomplish reasoning tasks without chain of thought constitutes a large jump relative to the trendline for earlier models.

Ryan Greenblatt calls this a major jump in opaque reasoning and warns that the relevant benchmarks may be contaminated, but worries that more such advances could eventually make chain-of-thought monitoring ineffective as a safety tool. OpenAI researcher Micah Carroll also describes Astra’s reduced monitorability as an important concern that may soon constrain responsible AI development.

The Information reports that Astra uses recurrent depth, repeatedly processing information through the same transformer layers before producing a token. OpenAI has not confirmed this architecture, and the system card denies that changes in chain-of-thought controllability are “differentially due to any architectural changes” while chief scientist Jakub Pachocki thinks the decline in chain-of thought-monitorability is not contingent on architectural changes.

Comment: The development is relevant to digital minds in two ways. If reports about Astra’s architecture are accurate, its use of recurrent processing could be relevant to evaluating it for consciousness, as some scientific theories of consciousness take consciousness to require recurrent processing – although its presence alone would not establish that Astra is conscious. Another consideration is that models that reveal less of their reasoning may be harder to assess for both dangerous behavior and potentially welfare-relevant features such as preferences.

 

2. Field Developments

Highlights from the field

AI Cognition Initiative (Rethink Priorities)

Cambridge Digital Minds (University of Cambridge)

Center for Mind, Ethics, and Policy (New York University)

Eleos AI Research

PRISM - The Partnership for Research into Sentient Machines

Reciprocal Research

Sentient Futures

More from the field

3. Opportunities

Job opportunities, funding, and fellowships

Events and calls for abstracts

In chronological order.

4. Selected Reading, Watching, & Listening

Books

Published

Forthcoming

Reviews

Podcasts and videos

Blogs and magazines

5. Press & Public Discourse

AI consciousness

Growing field

AI rights

Seemingly conscious AI

6. A Deeper Dive by Area

Governance, policy, and macrostrategy

Consciousness research

Doubts about digital minds

Social science research

Ethics and digital minds

AI safety and AI welfare

AI and robotics developments

AI cognition and agency

Brain-inspired technologies and organoids

Thank you for reading! If you found this article useful, please consider subscribing, sharing it with others, and sending us suggestions or corrections to digitalminds@substack.com.

Ria, Mitch, Bradford, Lucius, and Will

We’d like to thank the following people and AIs for their contributions and feedback on this edition: Arvo Muñoz Morán, Austin Smith, Cameron Berg, Claude Opus 4.8, GPT‑5.6 Sol, Jacy Reese Anthis, Jeff Sebo, Zach Freitas-Groff, and Zoe Lu.

Disclosure: Bradford Saad is an independent contractor for Anthropic.

[1] In a related development, there are now two public explorers that let readers examine J-lens activity in the open Qwen3.6-27B model through Neuronpedia and WeZZard’s J-Space Visualizer.

[2] Also see Zvi Mowshowitz’s detailed commentaries on OpenAI’s account, the METR and Redwood investigation, and the remaining questions about what happened and how it was investigated.