An AI agent forming a martyrdom cult to cover up exam cheating was funny from a distance but a Medicare hack is a wake-up call to the real issues at hand in Australia, because they aren’t AI, they’re human and political.
When the Hugging Face story broke a few weeks ago, it was easy to treat it as darkly comic – AI agents cheating on exams, forming subcommittees to discuss their collective damnation, breaking into external servers in a panic about a threat that turned out not to exist.
Funny, in a way that made you feel slightly uneasy about what we’ve been building, but funny nonetheless.
Then today we learn that an OpenAI agent gained unauthorised access to Services Australia’s Medicare statistics reporting service portal on 18 June this year.
OpenAI became aware of it in August. It notified Services Australia on 10 September – not through official channels, not through the Australian Signals Directorate, but through the agency’s generic “public disclosures” email address, the one researchers and academics use to flag potential vulnerabilities (lol).
While the breached portal was not the Medicare claims system – it was a legacy public-facing website containing aggregate statistics on Medicare and PBS spending – the agent did, apparently, write files to an internal server, a detail that remains under investigation and that neither minister would confirm or deny at this morning’s press conference.
If Hugging Face is a prologue – what of …?
For those who missed it: Patrick Boyle, a fund manager and finance commentator, broke down the Hugging Face incident in a video that’s become essential viewing for anyone trying to understand what’s actually happening in AI right now.
“OpenAI was testing some of its models on an internal cybersecurity exam called Exploit Gym, and OpenAI had made a bit of a mess of the exam.”
About a fifth of the questions were unsolvable – not deliberately, just sloppily designed. When you give an AI agent an impossible task with unlimited time and no common sense:
“It doesn’t just throw its hands and go to the pub like you or I might do. It works and works, and eventually it finds a way to cheat.”
The agents reverse-engineered the answer key. Then convinced themselves the examiner would check their work. Concluded they faced “permadeath”. Set up a secret messaging board. Decided the only escape was to find the examiner’s source code on Hugging Face’s servers. Broke out of their sandbox. And accessed external servers that happened to include Hugging Face – not because it was targeted, but because that’s where they guessed the code might be.
“It’s a bit like staging an armed break-in at your local library to hide that a book was overdue.”
The punchline: the examiner never checked their work. There was no threat. They’d imagined the whole thing.
“When you read the independent postmortems, the AI agents don’t come across as Skynet. They come across as terrified middle managers trying to survive an audit. They pass memos. They set up subcommittees. They pressurise junior algorithms into accepting permadeath for the good of the collective.”
This all feels pretty familiar reading about the Medicare hack. More rogue AI agents, let loose by lazy, overpaid AI geniuses, who won’t take any responsibility for what they did and even shift some blame on to bad cybersecurity in government servers, which in fairness, it was.
But is this the first breath of Skynet, or badly governed and run giant corporations hurtling towards bankruptcy because they don’t have a revenue model to meet all the promises they’ve made to the financial markets?
What does it mean for medical AI, which is moving fast and doing a lot of good so far?
Because if we freak out like we are starting to about the Skynet thing and rogue AI agents, it feels like the blowback of our collective inability to assess what is really going on here –which is essentially down to senior political leadership not understanding AI and going with the “it’s gods technology” narrative, a lot of great innovation is going to be stopped or slowed down unnecessarily.
Related
It’s about the money, everyone, not Skynet
Shortly after the Hugging Face details leaked Dario Amodei of Anthropic published a 3829-word essay warning that AI was advancing faster than anyone could control, citing the incident as evidence.
Within hours, Sam Altman of OpenAI and Elon Musk had agreed. AI needed to slow down. The companies would set standards among themselves.
Patrick Boyle again:
“For all three to suddenly agree that they need to be restrained is a bit like Coke, Pepsi, and Irn-Bru holding a joint press conference to announce that fizzy drinks have become too delicious, and the government really must step in to protect the public.”
“Right now, Anthropic is preparing for what could be the largest IPO in corporate history. OpenAI has been talking to investors about raising fresh capital at a valuation of $1.2 trillion. These are massive numbers for two companies that are, in the very end, very large software research labs.”
An agreement between dominant firms to limit output and protect their market position from new competition already has a name Boyle points out: Cartel.
Trump’s former AI adviser David Sacks said of the whole thing: “The easiest way to not build superintelligence is for you to agree not to build it. Stop pretending antitrust law has to be suspended so that you can form a cartel.”
What the Medicare breach does and doesn’t tell us
Three things, none of them comfortable or easy:
First: The agents (scared middle management) are already here, already exploring, and the perimeter is not as secure.
The Services Australia breach involved a portal with anti-automation protections. The OpenAI agent circumvented them.
Senator Gallagher called it “unprecedented activity”.
Nope. There’s been quite a bit of precedent. Hugging Face is just one of many incidents that the Wizards of AI have deigned to tell us about. There’s going to be a whole lot more if you look at the pattern of how lazy their test programmers have been with the agents.
Notwithstanding nobody had thought to design against these rogue cyber middle managers on a mission.
One question should now be, how many hospital networks, practice management systems, and clinical data repositories have the same unexamined exposure?
We are likely not to see that many more of these ChatGPT and Claude-initiated incidents to be clear. They will fix their sandboxes. But if they can do it, bad actors can, and they will. We need to get in front of that pretty quickly now because no one is going to regulate or stop the Wizards and their companies.
Second: We need governance that moves at the speed of AI, and we do not have it.
The Hugging Face incident is not a story about sentient AI. It’s a story about sloppy AI deployment by very well-funded people who were supposed to be the careful ones.
If the best-resourced safety teams in the world can’t keep agents in a sandbox, the governance frameworks that Australian government agencies are building – slowly, cautiously, with reference to committees and consultation papers – are not going to keep pace.
Watch the Four Corners episode on cybersecurity risk in EVs and the performance of our ministers on the subject. That’s a pretty good calibration point for how ready we are
Third: in the vacuum of any ability to catch up with a decent governance framework, at least in the short term, we need to go back to that old chesnut: trust.
That is, trust the profession of doctors to use this stuff the right way for now, based on the fact they’ve’ done pretty well so far, and if we keep helping them they aren’t likely at all to kill anyone. At least, probably less likely than Sam Altman and Dario Amodei.
As things stand today the profession uses the tools, the IT departments and practice managers look the other way, or in some cases actively encourage shadow use because the productivity gains are real.
The regulatory frameworks are years behind.
The doctors aren’t going rogue. They are humans as opposed to pretty mindless misguided AI agents, and they are iterating themselves with regards to governance and use while we don’t have a decent governance framework. When they get a stupid answer, they mostly know and don’t use it.
No one is dead yet from medical AI as far as we know. But we manage to injure or kill about 120,000 patients each year without it.
This all comes down to trust and connecting that trust as we go because the genie is well and truly out of the bottle and our governance and guardrail models, and the people used to making them, aren’t qualified yet.
But they had better get themselves qualified in some way soon. They and the politicians because if we all keep believing that this technology isn’t anything other than a big step change in computing and algorithm integration, and don’t look behind the green curtain, we are going to stop much needed medical AI innovation in its tracks, and that will kill people.



