hroelant's Quick Takes

By Harm Roelant @ 2026-09-09T13:25 (+1)


Harm Roelant @ 2026-09-11T14:39 (+1)

Anthropic's September report on misuse of its models poses interesting questions about AI safety. The report, detailing actions ranging from the use of Claude to set up a fake online dating profile farm to the model being employed by the Houthis to design software to guide missiles, was not produced under any legal obligation. Under the TFAIA in California and the EU AI Act, it is only mandated to confidentially disclose safety incidents to authorities, and the jury is still out on whether some of these cases would fall under the definition of 'serious incidents' (in the EU AI Act case) or 'critical safety incidents' (in the TFAIA case). Potential explanations for why the company still decided to produce such a report are not hard to imagine: better PR, internal employee pressure, and institutional inertia. However, these incentives are largely internal to the company and can change. If the board decides that these reports should be toned down to dissociate Anthropic from dangerous uses of AI (perhaps to prepare for an IPO), there is no reason why this would not take place. That would be a serious blow to the cause of AI safety at a time when the potential fallout from increasing model progress has still to be spelled out. Any loss of data on this front would mean one more threat vector left unexplored, and a part of society and the world left potentially unprepared.

hroelant @ 2026-09-09T11:50 (+1)

Within one week, the firm Calif created a zero-click hacking tool that could compromise WeChat accounts of targeted people according to recent NYT reporting. No matter how well intentioned such an exercise is from the start-up, and Calif does indicate that its tools was built experimentally to benefit cyberdefences, it seems impossible to prevent such technologies from being weaponised in the future by states, and for other states to interpret it as such (everyone can imagine how a tool specifically attacking a WeChat vulnerability must look to the Chinese government..).