Opinion: The debate over the government's aged care algorithm is missing the point


By Dr Andrew Rochford BMedSc MBBS (Hons)*
Friday, 21 August, 2026


Opinion: The debate over the government's aged care algorithm is missing the point

Two things are being said about the government’s besieged aged care funding algorithm. The Minister says it delivers predictability and consistency, and that it answers the Royal Commission into Aged Care Quality and Safety’s call for nationally consistent assessment. The Opposition says there is no evidence it is working. Both can be true at once. That is the problem.

Consistency is not correctness. If the standard itself is wrong, the tool applies that wrong standard to every person it assesses without ever varying, and an error applied uniformly stops being a mistake and becomes policy. The Royal Commission did not ask for a system that decides the same way every time. It asked for a system that decides well. Those are two different tests.

What neither side has put on the table is whether the algorithm reaches the conclusions a qualified assessor would reach about the same person. The government can show the tool applies its own standard consistently, because that is what software does. The Opposition is right that no evidence has been published, but that is what you get when the comparison has never been run.

I say this as a doctor who has used decision support tools to help make clinical assessments. The tools are often good. But they are not the judgement. Every clinician can name the moment a case stops fitting the formula because the context has changed. An algorithm cannot. Whether a 96-year-old woman living alone can manage is not purely a scoring exercise. It is a clinical and human judgement about what dignity requires for that person.

The science to settle it already exists

Human judgement is fallible too, which is why medicine stopped asking anyone to take our word for it. Accredited pathology laboratories are sent identical samples, test them blind, and are scored against the consensus of their qualified peers. Nobody re-checks every test. The laboratory itself is what gets measured, against a standard that is not a document on a shelf but the living, current judgement of the profession. That has run since 1968, in a profession performing billions of tests a year.

Drug trials rest on the same logic. Nobody tests a new medicine on every patient. A properly drawn sample, assessed blind and monitored over time, licenses a medicine for millions on the strength of a few thousand patients.

Sampling. Blinding. Measurement against professional consensus. Continuous monitoring. Four tools, each proven over decades, each protecting millions of people right now. Nobody has pointed them at an algorithm deciding who gets care. That is the whole gap. Not missing science. Unapplied science.

The Opposition is right that care decisions should be evidenced. I would go further: they must reflect professional judgement. But putting a human in the loop to review every decision does not scale, and the research consistently shows reviewers placed in front of an algorithm drift into agreeing with it. Automation bias is not a theory. A tired reviewer approving their two-hundredth recommendation is not oversight.

A third option: Above-the-Loop Governance

The debate centres on whether to keep the tool or scrap it. A third option exists, and aged care already holds the ingredients.

Clinical assessors are visiting these people and forming professional judgements. That work is already happening. It has been the standard for years, and it is a benchmark that moves as the profession’s thinking moves. On a properly drawn sample of cases, capture the assessor’s conclusion before the algorithm’s output is visible, so that neither anchors the other, then measure continuously how closely the two agree.

This is what I have spent the last year working towards, in collaboration with the University of Sydney’s Centre for AI, Trust and Governance. It’s called Above-the-Loop Governance and it measures whether AI decisions meet the professional standard qualified humans are held to, at scale, through a ‘living benchmark’ that is set by experts and updated in real time. If AI decisions begin to drift from the standard, its authority over those decisions can be withdrawn and control returned to human experts. It also identifies the decision classes where the algorithm does not meet the standard and should not be deciding alone.

Under in-the-loop review every case is bottlenecked at a human, so the cost of governance scales with decision volume. Under Above-the-Loop Governance, the population proceeds at volume while a statistical sample is measured against the human benchmark, so the cost scales with statistical sufficiency rather than with volume. Source: doi.org/10.2139/ssrn.6981538 (click image to enlarge)

A few hundred well-drawn cases speak reliably for every decision the tool makes, because confidence comes from how a sample is drawn rather than how big the population is. That gives aged care evidence of whether the tool decides the way qualified assessors would, warning when it drifts, and an auditable record the department can point to when someone asks for a review.

The tool does not need to be thrown out, and the tens of thousands of people waiting cannot afford the pause that scrapping it would cause. What it needs is confidence, and confidence in a system like this is either a finding or a claim. Right now it is a claim, asserted on one side, contested on the other, and worn down a little further with every review request. Measurement is what turns it into a finding: continuous, published evidence that the tool performs the task it has been given to the standard the profession expects, consistently. And a finding, unlike a claim, does not require anyone to take the department’s word for it, or mine.

So, the question worth putting to the department is simple. Has the algorithm’s judgement been measured against what qualified assessors would decide for the same people? Not when it was trained. After it was deployed. And continuously, for as long as it holds the authority to decide. If that measurement exists, publish it. If it does not, that is the finding.

*Dr Andrew Rochford BMedSc MBBS (Hons) is a medical doctor, healthcare executive and the author of the paper ‘Above-the-Loop Governance’, a methodology for measuring whether AI decisions meet the standard a qualified human professional would apply. Rochford’s paper has been reviewed by academics at the University of Sydney’s Centre for AI, Trust and Governance (CAITG). The statistical methodology has also been independently reviewed by Dr Bradley Rava, a lecturer in Business Analytics at the University of Sydney Business School. You can read the paper here.

Top image credit: iStock.com/O2O Creative

Related Articles

Top 3 benefits of aerobic and strength training for chronic disease prevention

Australian researchers followed more than 8000 Brisbane adults over nine years, identifying three...

A Day in the Life of a social worker specialising in FASD

Prue Walker is a social worker specialising in Fetal Alcohol Spectrum Disorder (FASD), a lifelong...

Home care safeguards when disasters hit

Home care presents unique challenges for emergency planning. An in-home care provider's CEO...


  • All content Copyright © 2026 Westwick-Farrow Pty Ltd