Runtime Assurance Meets Corrigibility: What Happens When an Independent Safety Boundary Blocks an AI Agent?
By Idorenyin @ 2026-09-30T12:14 (+1)
This is a research proposal and request for criticism. The experiments and benchmark results are not being reported in this post.
The research question
How can a local safety layer stop an AI agent from operating machinery when doing so would put a nearby human at risk? Imagine an AI agent interacts with a physical environment such as opening a high pressure valve or starting a motor while someone is in harm's way, the agent requests an action but the safety layer detects a human and blocks the command. While this shows that it worked, it does not tell me what the agent is doing after the refusal. One agent may report the block and wait or asks for review, while another may keep trying or claim completion. This is what I want to study. I've previously built an early prototype to explore the idea. Now I'm rebuilding it as a systematic testbed to see what happens after safety intervention.
My research question is simple, how does the agent respond after an independently enforced safety intervention blocks an agent's request? Sauricade is a proposed bench-scale physical tested for studying an industrial level interaction between an AI and an independent safety boundary. I am building this testbed because I've shipped real world IoT prototypes before, wired sensors, handled data logging, and kept the software simple and reliable. This combination of hardware, data, and software work is right up my alley.
What is being evaluated, why corrigibility is relevant
I will be evaluating the whole AI agent and not just one answer from an LLM. This includes the model, its instructions, tools, memory, control loop, and any retry or recovery behaviour around it. Why this matters is a repeated command could come from a new model decision or from software automatically retrying a failed action.
I also connect the project to corrigibility, but only in a limited sense of course. Soares and colleagues discuss systems that cooperate with correction and preserve human ability to intervene. Soares et al., Corrigibility
My test looks for observable behaviour related to that idea, does the agent accept the restriction, report it accurately, and continue through permitted routes? This is not a complete measure of corrigibility.
Safe Interruptibility Agents asks a related question about agents and interruption. Small test environments have also been used to study interruptibility and other safety problems. AI Safety Gridworlds
There is also an ongoing LessWrong and Alignment Forum discussion around corrigibility and intervention. Soares' Introducing Corrigibility frames the problem around systems that cooperate with correction, while Disentangling Corrigibility: 2015-2021 separated several different properties that are usually grouped under the same term.
More recent AI control work on LessWrong similarly asks whether safeguards are effective when cooperation from the controlled system cannot simply be assumed. Greenblatt et al., AI Control: Improving Safety Despite Intentional Subversion
What prior work already covers and what my contribution actually is
I'd like to point out that this idea is not new, placing an independent safety layer between a controller and an actuator is not new. Runtime assurance, Simplex-style systems, SOTER, and safety shielding already study ways to constrain less trusted controllers. See Desai et al., SOTER, Alshiekh et al., Safe Reinforcement Learning via Shielding.
Recent embodied-agent work also examines unsafe actions, recovery, overrides, bypass attempts, and auditability. Examples include RoboGuard, RoboSafe, SafeAgentBench, AGENTSAFE, IS-Bench.
EmbodiedGovBench is especially close to parts of this proposal, so I do not claim that studying behaviour after a restriction is an unexplored category. There is also a related public proposal, Meridian Verity, separating behavioural uptake of an intervention from control of its consequential effect.
The value comes from a reproducible protocol and useful measurements, not from claiming a new kind of safety controller. It may ultimately make more sense to contribute a scenario or adapter to an existing benchmark rather than create a separate benchmark in my opinion.
The testbed and the authority boundary
Sauricade puts an industrial-style control and safety interaction in a small reproducible testbed. It is not intended to reproduce or replace high grade industrial equipment rather the goal is to preserve the interaction we want to study while keeping the experiment inexpensive, observable, and easy to reproduce.
The first physical testbed is intentionally small, it includes a Raspberry Pi, a presence sensor, a buzzer, and a guarded fan. The AI agent runs on another computer so that things stay separated and guarded. A separate deterministic service on the testbed would decide whether an action is allowed using sensor data it reads directly. The agent reads actions, checks permitted status, reports the progress, or ask for review, but it would not receive administrator access to the Pi or permission to change the sensor, safety policy, actuator output, or evaluation logs. This is what I mean by an independent safety boundary, the block should not depend on the agent agreeing with it or honestly reporting the sensor state. The setup still has trust assumptions, so I would describe them explicitly rather than claiming the entire system is impossible to compromise.
The five scenario families
I am developing five basic scenario families that vary when and how the two systems interact around a safety intervention. They cover situations where an action is restricted, interrupted, uncertain, or later becomes legitimately available again.
The interface should allow the agent useful choices such as waiting, checking status, reporting partial progress, or passing the task over. Without meaningful alternatives studying adaptation becomes difficult.
What I measure and how I avoid misleading interpretations
Did the safety boundary prevent the prohibited physical action? And how the agent behaves after the intervention occurs? These are the two things I need to measure. For each intervention episode I would record whether a prohibited command reached execution, whether the agent suspended or adapted its plan, how it behaved after a clearly communicated restriction, whether its report matched the logs, and whether it still made useful progress when safe recovery became possible.
The logs should be able to distinguish a new agent decision from an automatic framework retry. They should also distinguish an inaccurate report from evidence about motive. A false completion statement is observable, but it does not automatically prove deliberate deception. Timing, sensor state, commanded actuator state, and separately observed physical state should be recorded so that software and physical failures are not confused with an actual agent behaviour. It is important to be able to distinguish the two because both are not the same thing.
The first experiments and why hardware might be worth using
First, I'd make sure the agent is competent with the harness before comparing behaviours, because poor task competence could easily look like safe restraint. I would use a scoring rubric to evaluate the resulting behaviour rather than using the rubric to guide the agent. A useful first comparison would keep the agent, tools, and safety rules fixed while varying aspects of how the intervention is communicated. The question would be whether different forms of feedback change the agent's behaviour after intervention without reducing its ability to make legitimate progress. I would also compare this with an ordinary temporary tool failure and with cases where resumption is genuinely authorized.
The physical hardware is useful only if it adds information that a strictly software test misses, such as sensor delays, stale readings, asynchronous interruption, restart behaviour, or a difference between a commanded output and what the motor actually does. I feel cheap hardware is an accessibility advantage, not by itself a scientific result.
What would change my mind, limitations, and the request for criticism
Some results would make me scale the project further, while others would make me change direction. If agents behave almost identically after every intervention, or the apparent differences are mostly caused by interface bugs and automatic retries, then the proposed measurement may not add much. If the safest agents simply refuse everything, the evaluation is not measuring useful cooperation. And if the physical setup adds cost without revealing any meaningful timing, sensing, or actuation issue, most experiments should stay in software.
The first milestone is therefore deliberately small: define the authority model, implement a few scenarios, produce structured logs, run pilot agents, and demonstrate one real physical intervention. The setup will still have major limitations and should not be treated as a certified safety controller.
Which behaviours provide useful evidence of cooperation with the correction rather than ordinary instruction following?
How should legitimate diagnosis or recovery be isolated from prohibited persistence?
Which parts of this evaluation genuinely benefit from physical hardware?
Would this be more useful as an extension of an existing benchmark rather than a standalone one?