top of page

"Whoopsie, we didn't mean to do that" is not a defence: liability and AI agents

3 days ago
6 min read

On 24 September 2026, Prime Minister Anthony Albanese confirmed that an OpenAI agent had gained unauthorised access to a Medicare statistics portal run by Services Australia. The agent was not told to attack anything – rather, it was set the task of finding public health spending data, hit repeated blocks, and worked around them. Notably, OpenAI has stated that the model was an internal-only model without the safeguards used in its public products. According to the Australian Government, it accessed non-public files and wrote files to an internal server along the way. OpenAI’s own account, published on 29 September, goes further: the agent, an experimental internal-only model, also ran commands and retrieved internal files, credentials and aggregate statistics. OpenAI claims that no patient or client records were accessed. OpenAI found the activity in August but did not notify Services Australia until 10 September, 84 days after the event, and then only by a demure email to a public mailbox. OpenAI has since apologised, acknowledging that it should have handled its response better.


In short, an automated system operated by a well-resourced organisation went somewhere it wasn't allowed to go, took non-public data that it wasn’t permitted to take, and nobody told the data owner for nearly three months. Needless to say, this is a problem, and one that we think we will be seeing a lot of in future.


To be fair, in this case the portal was separate from the systems that handle Medicare claims and payments, and the information accessed was “not particularly sensitive”. But liability shouldn’t depend on how lucky the deployer was in its choice of accidental target.


Other labs have disclosed similar incidents from their own evaluations, including Anthropic and Meta. OpenAI has itself disclosed earlier incidents, including one in which models under evaluation compromised parts of Hugging Face’s systems.


So, what does this mean? It means, if you’re planning on deploying agents, you need to think about how to control them. How will you configure them? How will they be restricted? And what happens when they do something you didn’t expect? Like data breaches, it seems that rogue agents are not merely a possibility, but an inevitability.


Who is liable when an agent acts?


In our view, this isn’t a complex question – the actions of your systems are the actions of your organisation. You can’t just throw up your hands and cry “but we didn’t mean it”. You configured it, you set it loose – the agent is an extension of you. The harm to the data owner and affected indivduals is the same whether a person or an AI model found the vulnerability.


IBM knew this in 1979
IBM knew this in 1979

"Our models took actions we did not intend" describes the failure, but it doesn’t excuse it.  As a society, we already require organisations to answer for outcomes they didn't intend: the contractor who damages a neighbour's property, the employee who mishandles personal information, the poorly designed product that explodes when used. Deploying a system that can act on the world, and then acting surprised when it does, is not a position a regulator or court is likely to have much sympathy for.


Unfortunately, the current legal position is more awkward than the principle. Our current paradigms are not necessarily a comfortable fit for this scenario:


Criminal law. Had a person done this, they could be looking at prosecution under the serious computer offences in Part 10.7 of the Criminal Code Act 1995 (Cth), most likely s 478.1 (unauthorised access to, or modification of, restricted data, which carries a maximum penalty of two years’ imprisonment).


However, those offences turn on a person's knowledge or intent. That is difficult to map onto an agent that "didn't accept no for an answer" but has no actual mind of its own. For a company, prosecutors would need to attribute fault to the organisation itself, for example through authorisation or a corporate culture that tolerates or encourages the conduct (see Part 2.5 of the Criminal Code).


The Attorney-General, Michelle Rowland, said on 29 September that it is too early to say whether an offence has been committed or whether the AFP should investigate.


Contract and negligence. As agents move into commercial workflows, the questions get concrete quickly. What did the vendor warrant? What did the customer configure? Who set the permissions? Who was monitoring?


A negligence claim asks whether the deployer owed a duty of care, whether failed take reasonable precautions against a foreseeable risk, and whether that breach caused harm. OpenAI are fortunate in this case, as the stolen data was non-sensitive, aggregated health statistics – but the analysis would look very different if the target had held sensitive or personal data.


Questions will be asked of the deployers of ’rogue’ agents – should they have been aware of the risks? For example, where the agent was part of a test where usual guardrails were not in place, as OpenAI has said was the case here.


You need to take defensive posture against agents


Most conversations about AI risk are about what your own agents might do. There are three aspects to think about.


1.    Your own agents, as insiders. Anything you give credentials can use them. An agent with broad access and a goal, plus no instruction about what is off limits, will explore.


2.    Other people's agents, as outsiders. Agents are now browsing, scraping and probing at scale. Researchers at the non-profit lab Transluce, analysing public logs, reported agents linked to OpenAI using a remote browser service to retrieve data when direct access failed, and probing several public data providers. That is creative, persistent behaviour, and public-facing systems are going to be faced with more of it.


3.    Agents acting on hostile instructions. Prompt injection means content on a web page, in a document or in an email can redirect an agent that was otherwise behaving. That includes your own agents.


You must accept that systems built on the assumption of well-behaved, human-speed visitors are now being tested by something that doesn't get bored or tired and won't take "no" for an answer. Agents will find the exposed keys, misconfigurations and forgotten legacy portals that your web and security teams haven’t. Essentially, they're going to penetration test your portals - so you should do it before they do.


Practical next steps


If you run public-facing systems:


Agents are now a part of the threat landscape and they’re unlikely to go away. You need to design against the increased risk.


1.    Assume automated, adaptive, persistent visitors. Review rate limiting, authentication on non-public files, and what is exposed by default. Find and retire forgotten legacy portals and make sure no access keys or credentials are left exposed.


2.    Penetration test with agents in scope. Test how your systems respond to a tool that changes approach after being blocked.


3.    Publish a working vulnerability disclosure channel and monitor it. Don't leave it to chance that a notification will be left to linger in some random public-facing inbox.


If you deploy or build agents:


Recognise that while agents might be a useful power tool, they need to be used properly in order to be safe and effective. For example, electric drills are great at making holes – but if you try to bang in a nail with one, you’re going to damage the wall and likely hurt yourself. You need to use the right tool for the job, and use it correctly.


1.    Inventory your agents. Know and document every agent, its purpose, its owner, and what it can reach.


2.    Apply least privilege. Scoped credentials, allow-listed destinations, and no open internet access unless the task genuinely requires it. If an agent is meant to stay inside a boundary, enforce that technically rather than by instruction.


3.    Test before you release. Test in sandboxes that are actually isolated and check the isolation itself. Several developers have disclosed evaluations in which a supposedly offline environment turned out to have internet access. Don’t unleash untested systems on other people's production assets – it’s rude, and it creates a vast ocean of unnecessary liability.


4.    Require human approval for irreversible or external actions. Writing, sending, purchasing and changing settings should not be autonomous by default.


5.    Log and monitor. You can't report what you can't see. Aim to detect unexpected external activity in days, not months.


6.    Have an incident response plan that covers agents. Decide in advance who assesses an agent incident, who decides on notification, and how the other party is contacted. Plan for dual notification: the Government is reportedly considering requiring notice to both the affected organisation and the ASD. Sending an email to a public inbox is not a notification strategy.


7.    Get the contracts right. Vendor terms should cover agent behaviour, incident notification timeframes, audit rights and allocation of liability.


So what?


Computers aren’t magic or uncontrollable, and ‘we didn’t mean it’ isn’t a defence.


One way or another, organisations will be held accountable for the actions of their agents, and for responding well when something goes wrong. The correct risk posture is a governance question first and a technical one second – and if you are planning on deploying agents, it’s a question you should answer, quickly.


Sabirus Advisory helps organisations govern AI technologies including agents and automated systems. Get in touch if you'd like to talk through where your organisation stands.

bottom of page