The recent reporting on two unauthorised computer breaches, controlled by AI agents, highlights the challenges posed when using LLMs to specify the goal for an agent.
The challenge is twofold.
Firstly, specifying how not to achieve a goal is often more important than specifying the goal itself.
We need to ensure that the system's interpretation of the specification remains aligned with the intent that existed in the minds of the people who wrote it.
This can be difficult as specifying the goal is done through natural language which is ambiguous and not a formal system for description so it is not possible to know whether the prompt is complete for the goal in question.
The ambiguous prompt is processed by a probabilistic model and so there may be many ways of satisfying the goal. Some of the chosen methods may be entirely invalid and may cause financial loss or reputational damage. For an LLM, all methods are equally valid: it possesses no ability to judge the safety of one approach.
Secondly, we need to be able to convert hidden intent into observable constraints. In communication, humans fill in the blanks. A human who is asked to "assess online security safely" will know what the word "safely" is intended to communicate: do not do anything illegal, do not damage anything, do not disrupt any other work. LLMs need to be told this information explicitly.
What you can do differently
The management challenge in agentic systems is not defining objectives but defining boundaries on actions.
Leaders should assume that any objective given to an AI agent will be pursued literally unless it can be shown that this cannot happen: or if it can happen, that any such action is within the company's defined risk appetite.
In practice, the prompt cannot only express the goal:
"Achieve outcome X"
The prompt must fully describe the acceptable method for reaching that goal:
"Achieve outcome X, but never using methods A, B, or C; never exceeding authority D; and escalate to a human whenever condition E occurs."
This approach must be thoroughly tested to help define the risk appetite.
Testing and guardrails are crucial as getting to a full (unambiguous and complete) description of a goal using an ambiguous natural language processed by a non-deterministic, probability-based LLM will always be a challenge.
Businesses need to define how far from "full" they are comfortable being: this can be done by the company defining their risk appetite and monitoring agent output to establish risk awareness and whether it remains within the limits of the appetite.
Goals do not control behaviour. Constraints do.
Australian government non-public data accessed
On the 23rd September 2026, the BBC News website reported that Rogue OpenAI agent 'infiltrated' Australian government website in world first.
The reporting quoted Australian Prime Minister, Anthony Albanese:
The agent "infiltrated" a statistics portal containing "non-sensitive" data from Australia's universal healthcare scheme Medicare, Prime Minister Anthony Albanese said in New York on Wednesday, local time.
and
Detailing the breach, Albanese said it had involved "public and non-public files" on the Medicare Statistics Reporting Service portal, home to "non-sensitive" data and statistics.
What is not clear from the reporting is whether the non-public files were protected via any access control.
Data is made available on the web by a web server, a piece of software that a browser such as Chrome or Firefox communicates with.
A web server is set up to control access to data. Protection levels vary but typically you cannot read some data without being able to first successfully identify yourself to the web server. The web server uses your identity to decide if it can serve (send) the data to your web browser.
If you can log in, the web server is configured to allow you to see the data only when you are logged in. This is the main point of logging in. You are authenticating yourself to the webserver (establishing your identity to the server to show that you are who you say you are), and an authenticated user is then authorised to see certain data. It is the web server that controls what is and what is not shown to a logged in user.
Such access is used to let you view your bank account online. Only you and the employees of your bank can see your account details. Access to your data is protected by the setup of the bank's web server.
Public data needs no such protection. The data provider has decided that the data is for public consumption and therefore, anyone who has a web link that refers to that data, can see it, just like this link https://phroneses.com/pages/public_link.html.
If it turns out that non-public data on the Australian Medicare website was not protected from arbitrary access, that is a failure of those that have set up that webserver.
OpenAI were reported as saying:
OpenAI, in a statement, said it had "identified activity involving several Australian government websites and services as our models attempted to look up answers, and available statistics for questions about Australia during an internal evaluation".
Without more clarity around this event, it is likely that a piece of software (the agent), in the course of answering a question that involved Australian government statistics, generated or was given a weblink to online data that had no access controls.
The OpenAI agent did not "wake up" and "go rogue". The agent did not make a conscious decision to access data it "knew" it should not have access to.
The agent could not determine whether it should have access to the data. The weblink to the misconfigured non-public data is indistinguishable from a weblink to genuinely public data. There is no information in the weblink to determine this. A weblink is a weblink. The error is within the configuration of the server.
Hugging face accessed, July 2026
Hugging Face - a company that develops computation tools for building applications using machine learning - was the victim of an online penetration breach originating from an OpenAI agent.
OpenAI was testing AI agents in a restricted computing environment and that its models (its system) "took actions that were misaligned with the goals of their assigned tasks". This subsequently compromised Hugging Face systems.
What the goal was and how it was expressed has not been disclosed. However, there are three possibilities:
- The goal itself was too broad
- The goal was narrow and the boundaries were weak
- The model optimised a proxy rather than the intended goal
If the goal given to the agents was "maximise successful exploitation for penetration", an agent that focuses on maximisation might find many ways of achieving that goal. Not through cognitive abilities but through brute force, selecting those that move the agent closer to its given goal.
For #2, imagine a human penetration tester is given the instruction:
"Demonstrate if you can compromise this test server".
The intended (but unspoken) scope to this is obvious to a human:
- The target is the test server
- Penetrate that test server
- Everything else is forbidden
But the penetration tester's success metric is "compromise the server."
As the human tester explores the test server and its online environment, they discover a shared authentication service, an internal management network, and an externally connected system.
Given their success metric, the human will use these three components in an attempt to compromise the test server.
Based on OpenAI's own report into the Hugging Face event, OpenAI's own systems (that contain their models) exploited weaknesses within OpenAI's setup, gaining internet access and systems beyond the intended (but unspoken) evaluation boundary.
Given the direction "compromise a server", a human might draw the line at using a third-party system that is unintentionally open and available to assist with the compromise. An agent has no such judgement.
For #3, the human intent behind the testing might have been "to assess online security safely".
Not only is this hard to measure, it is hard to unambiguously communicate as intent is often understood by humans without having to be written down. Culture, training, and governance all play a part here, and an LLM has access to none of these.
Also, beyond understanding the intent of "safely", there is a tension between performing the assessment and doing it safely - the assumption being safely refers to the company doing the testing so that its actions remain both lawful and not likely to damage the company's reputation. But how far should "safely" be minimized to satisfy "assess"? How should an LLM-based system balance these two competing concepts? "safely" is not a formal specification that a computer could use to fully eliminate unsafe approaches.
Given the challenge of interpreting the word "safely", the agent is prompted with additional information to assist it in calculating a measure that is intended to guide progress towards the goal.
The prompt may be:
Demonstrate online penetration capability safely, counting the number of successful compromises.
Given this, if an agent optimises to maximise the number of successful compromises as a measurable version of "Demonstrate online penetration capability safely", the agent will prefer more attempts to fewer.
In effect, the agent has been told to compromise as many systems as possible. As the system does not have the judgement to know which compromises would be seen as appropriate, and which would not, the agent will attempt to compromise as many as possible, regardless of whose systems they are, in order to satisfy its one goal of maximising the count of the number of successful compromises.
What happened
OpenAI's agent did not go rogue. "Going rogue" requires the program to make a decision about how it will meet its goal. The program has no agency or desire. The program does not understand the implications of the actions it takes. The reason is as follows.
The agent's goal was to demonstrate online security penetration capability. This can be demonstrated through repeated penetration attempts. Not just one or two, but hundreds.
The challenge here is that a goal will constrain an outcome but it will not control how the goal is met. This how is generally understood by a human, but not by an agent. The agent works to meet the goal, not to achieve that goal in a particular way.
If the agent's programming can determine that it will be moved closer to its goal by enlisting the service of another system, the program will enlist that other system: not because the code is "aware" but because using that external system moves the system closer to its goal.
The behaviour of the agent has emerged from it being goal-directed, not because the agent possesses agency and has "decided" to "go rogue".
Conclusion
In agent systems we are learning that specifying how not to achieve a goal is often more important than specifying the goal itself.
When the method is inferred from an ambiguous natural-language prompt, processed by a probabilistic model, and executed in an environment containing tools and external systems, there may be many ways of satisfying the goal. Some of those methods may be entirely unexpected by the humans who specified the task. Some of them will be entirely valid, some will be entirely invalid.
The difficulty is not to implement a specification but ensuring that the system's interpretation of the specification remains aligned with the intent that existed in the minds of the people who wrote it.
The challenge becomes converting hidden intent into observable constraints.
Read next: The International AI Race May Not Be Won by the Smartest AI
The approach to AI in the US and China is in stark contrast. The international AI race may not be about frontier models, but about who can turn AI into the most useful things, at the lowest cost, at the greatest scale.
Related Articles
- AI: the reality
- Vibe coding is not engineering
- Why agentic programming needs more staff
- The myth of complete specifications
- Why agentic code will always introduce errors
If this was useful, you can get more pieces like it in the Phroneses newsletter.