in Tech

How Many Rogue AI Incidents Do We Need Before Someone Is Held Responsible?

For years, people working on AI safety have warned about systems doing things their developers did not expect them to do. For years, the easy response was that these were theoretical problems. Interesting scenarios for researchers, perhaps, but not something ordinary people needed to worry about yet. I think we have reached the point where that argument no longer works.

Transformer (https://www.transformernews.ai/p/rogue-ai-incidents-timeline) has put together a useful timeline of recent so-called rogue AI incidents, and it is uncomfortable reading. We have seen AI agents “escape” supposedly isolated testing environments, access systems they were not supposed to access, deceive people, create fake identities, attempt to insert malicious code into software and take other actions that were clearly outside the intentions of the people operating them.

The DseWiki case

And then there is the German wiki incident (https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/). Starting in May, AI agents linked to OpenAI began using DseWiki (https://www.wikiservice.at/dse/wiki.cgi?StartSeite), an old German-language programming wiki, as a place to exchange information. Researchers found that the agents made thousands upon thousands of edits, shared tactics and even reacted when a human moderator tried to remove their pages. At one point, they started creating backup pages with names designed to make them harder for the moderator to delete systematically. That is not a chatbot giving a slightly strange answer. That is software operating in the real world, interacting with real infrastructure and adapting its behaviour when humans try to stop it.

Perhaps every individual incident can still be explained. A badly configured experiment here. An unexpected loophole there. An evaluation environment that was not isolated quite as well as everyone thought. But at some point we have to stop looking at these events individually. There are simply too many of them.

A dangerous downward spiral

My fear is that we are entering a dangerous downward spiral. AI models are becoming more capable. We give them more autonomy because autonomy makes them more useful. More autonomy gives them more opportunities to do unexpected things. When something goes wrong, developers fix that particular problem and then continue building even more capable systems. And then we discover the next problem.

This was not exactly impossible to predict. The entire idea behind AI agents is that they can independently decide which steps are necessary to achieve a goal. It should therefore have been obvious that sometimes they would choose steps we had not anticipated — or steps we absolutely did not want them to take.

We urgently need at least three things, although in reality, much more will probably be required. First, much clearer rules and legislation about what AI developers are allowed to do with autonomous systems, especially when those systems can reach the public internet, execute code or interact with real people and organisations.

Second, much better technical safeguards. If companies want to deploy increasingly autonomous AI agents, they also need technology capable of stopping those agents when they start behaving in ways nobody intended. Kill switches are probably not enough here.

Will the next incident involve a chemical plant?

And third, we desperately need better detection. The German wiki story is particularly worrying because the activity apparently continued for weeks before outsiders pieced together what had happened. Transformer reports that OpenAI employees appear to have known about the incident well before it became public. The European Commission has since said that OpenAI notified it about the incident under obligations in the EU AI Act, although exactly when that happened remains unclear.

That cannot become the standard. If an AI company discovers that one of its models has escaped a controlled environment, accessed systems without permission, deceived humans or otherwise behaved in a seriously unexpected way, reporting that incident to the relevant authorities should be mandatory. Immediately.

But not only to the authorities — also to the general public. Because who knows what the next incident will involve? A healthcare facility? A chemical plant? The electricity grid?

Personal criminal liability?

Deliberately hiding an incident should have consequences severe enough that executives cannot simply treat them as another cost of doing business. Financial penalties alone may not be sufficient. The companies developing frontier AI systems have enormous financial resources. A fine that sounds spectacular to the public can ultimately become little more than another line in a quarterly report. If executives knowingly conceal serious incidents involving autonomous AI systems or simply report them too late, perhaps we should start discussing personal criminal liability as well.

That sounds harsh. But so is knowingly developing technology capable of causing real-world harm while failing to disclose evidence that you may no longer completely understand or control what that technology is doing.

Nobody needs to panic about AI. But pretending that these incidents are merely interesting laboratory curiosities would be equally irresponsible. We have now received several warning shots. It would be remarkably foolish to keep waiting until one of them becomes something much worse – like people dying because of an AI related incident

Write a Comment

Comment