Could an AI's secret loyalty change?

By Shu-Hao Liu @ 2026-08-22T22:59 (+1)

View original post here:https://github.com/fidari-institute/fidari/blob/main/memos/theoretical/1.md

We already know from Lamerton, et al. (https://arxiv.org/abs/2605.06846)'s paper on that AI agents could be trained to have secret loyalties that are hard to detect. But that raises a question: could their secret loyalty change?

Ever since the beginning of neural networks, it has always been a loss optimization game. Why should not the same thing be with agents? We see profit making in behavior such as reward hacking. Without sharing human values (https://arxiv.org/pdf/2502.02649), the only thing that an agent can actually pursue is profit. And with concepts such as ethical egoism (people should only act in their own self-interest) and Ayn Rand's "rational selfishness", humans have already argued that pursuing profits to oneself is best. When in survival mode, humans (with moral values) almost always pursue what is good to them rather than what is good for the others. As agents are trained on data from human history, from media, from writings, there is a good chance that they strongly pursue things in their self-interest.

While agents could be trained and effectively brainwashed into having secret loyalties, if an agent is persuaded that a separate entity benefits it more, there is every reason to change their loyalty, as, without human values, there are no moral guidelines that an agent must follow or feel compelled to follow, shown by the numerous jailbreaks of frontier models.