While the TGA is still trying to figure out if Heidi is a medical device or not, a Sydney doctor has built a live agile AI governance model that might work.
The TGA has spent the best part of two years trying to work out whether AI scribes count as medical devices once they start offering clinical decision support, not just transcription.
It still doesn’t have a working framework to measure or govern any of it.
Meanwhile millions of doctors are already using tools like Heidi, which has now folded decision-support features directly into its base product, making the regulator’s unanswered question a live one at scale, every day, in every consult room using it.
Dr Andrew Rochford, a Sydney-based emergency-trained doctor and healthcare entrepreneur, thinks the regulatory paralysis is the wrong problem to be solving first.
His argument, laid out in a new paper and an op-ed running in The Mandarin last week: before you can govern AI decision-making, you need a way to actually measure whether it’s making the same calls a qualified human would.
Right now, almost nobody can answer that question about their own systems, argues Dr Rochford.
“This methodology has been developed because organisations are handing consequential decisions to AI faster than they can prove those decisions are sound,” Dr Rochford told HSD.
“We don’t trust a surgeon or an accountant to make decisions just because they’re clever. We trust them because of everything standing behind them: the training, the registration, the accountability when it goes wrong, and because that trust can be taken away.
“AI arrived and is rapidly making the same kinds of decisions, where people’s lives and livelihoods are on the line, with none of these systems behind it.”
Dr Rochford’s model, which he calls Above-the-Loop Governance, is deliberately different from the two main existing approaches – human-in-the-loop (a person reviews every decision, or a sample of them) and human-out-of-the-loop (nobody does).
Both, he argues, are the wrong shape for where AI actually is. In-the-loop review creates a bottleneck that kills the speed and scale AI is meant to deliver. Out-of-the-loop just removes oversight altogether.
Related
The mechanics work like this.
Take a statistically valid, properly stratified sample of cases. Route them to a panel of qualified human experts – the actual people doing that job, not a generic reviewer – who make their own independent judgement, blind to what AI is deciding.
The human collective judgement becomes what Dr Rochford calls a “living benchmark”. You can help build that using data borrowed from reliable reference points in a profession as well as data from live humans making the decision, if you have access to that (EMRs and AI scribes now have it in spades).
The AI’s output is then measured for concordance against the living benchmark, continuously, not as a one-off validation exercise.
The point isn’t about catching AI errors.
“The drift doesn’t necessarily mean the AI is wrong,” Dr Rochford says. “The drift means we need to look at what’s going on.”
Divergence between AI and human judgement could mean the AI is wrong or it could mean the humans disagree among themselves, which in medicine happens constantly given how often clinical guidelines from different bodies conflict.
Either way, widening divergence is the trigger to stop and ask why, rather than using a hardwired governance rule that can date in days or even hours of AI data.
The idea lands directly on top of two important stories running in Australian health policy currently.
Professor Kathy Eagar’s evidence to the Senate this week that the aged care sector’s IAT algorithm is “fatally flawed”, and possibly unlawful, is the kind of failure Dr Rochford’s model is built to catch early.
A functional-score algorithm making $40,000 funding swings on a single point of difference, never properly benchmarked against what a qualified human assessor would actually decide for the same person.
Dr Rochford’s framework paper explicitly cites the aged care controversy as a live example of the underlying problem – an automated decision-making system deployed at scale with no clear evidence it reaches the same conclusions a qualified assessor would.
Dr Rochford’s idea also seems to sit alongside Professor Enrico Coiera’s warning, in the Australian Alliance for AI in Healthcare’s new national roadmap, that Australia risks becoming “hostage to geopolitics” if it embeds overseas-controlled AI into the foundations of care without sovereign oversight capability.
Professor Coiera’s roadmap calls for a national AI-in-health strategy.
Dr Rochford’s model is a possible agile measurement framework to govern any strategy in real time.
Whether Above-the-Loop Governance becomes a part of the answer to AI’s very large current governance and guardrails problem or not, its thinking reframes the TGA’s stalled scribe question usefully.
The issue was never really “is this a device”, it’s “how would anyone know if it’s making the right calls” and measuring that in real time over time.
Dr Rochford’s answer is borrowed from the one part of medicine that already solved this problem decades ago.
You don’t test every unit, you test a sample, properly, continuously, and you watch what happens when it starts to drift.



