A copilot is only useful if you know how to fly

13 minute read


If we never develop the skill of asking the right question, how will we know whether the answer we have been handed is sound?


This week I started as Associate Professor in Paediatrics and Child Health at Western Sydney University, with a focus on digital health and AI.

Western Sydney is one of the most culturally diverse parts of the country and one of its fastest growing. Many of our students come from these communities, and most will go on to work in them. Teaching the next generation of clinicians here is a privilege, and the consequences of getting it wrong sit well outside the lecture theatre.

Then I sat down with the curriculum. I have taught and worked on digital health and AI at UNSW, at ANU, and at the UCL Institute of Child Health in London. I have also spent years on the other side of it, building a health technology company, and more recently ConsultPilot AI, our document intelligence tool. I assumed all of that would give me a running start.

What I found instead was how quickly parts of it have dated. Two papers published this year made that plain, and unsettled a lot of what I thought I knew about what we should be teaching, and in what order.

Never-skilling

Everything we are taught as clinicians rests on the same foundation, whether we call it clinical reasoning, critical thinking, or plain analytical thought. We spend years building a working model of what normal looks like, what abnormal looks like, and how to tell the difference when the picture is incomplete, which it usually is. That model is what lets us ask a useful question in the first place.

If you don’t know the baseline, you don’t know what to ask. And if you don’t know what to ask, you have no way of judging the answer that comes back.

A Nature Medicine Perspective published in June, led by a team at Duke-NUS, gives a name to what happens when that foundation never forms:

Deskilling is when an experienced clinician relies on AI and an established skill erodes through disuse. The foundation remains; the practice has lapsed.

Mis-skilling is when a trainee accepts an incorrect or biased AI output and internalises the wrong lesson from it.

Never-skilling is different, and it is the one I find most concerning. It is not the loss of a skill. It is the failure to build it at all, because AI did the reasoning during the years when the reasoning was the point of the exercise.

The learning theory behind it is well established. We learn by working at a problem, getting it wrong, and understanding why. Remove the difficulty and you may remove the learning, even when the answer that arrives is correct. The authors’ concern is not AI itself but the sequence. Their line about it puts it well: a copilot is only useful if you have first learned to be a pilot.

The authors are careful to say that never-skilling remains a hypothesis, that the longitudinal studies have not been done, and that direct evidence in clinical trainees does not yet exist. While all of that is true, I am not sure we need to wait for a cohort study to accept what is already visible in front of us.

I can see it in my own practice. I reach for AI evidence tools far more quickly than I did two years ago, and considerably more often. I will use OpenEvidence to confirm a management plan I have already formed, and sometimes, if I am honest, to ask for one. I don’t think that has made me a worse clinician. But I am a specialist. I have done the training. I know what I am looking for, I know which questions are worth asking, and I know when something doesn’t smell right. The tool is sitting on top of a foundation that took years to build.

The trainees don’t have that yet. The medical students certainly don’t. They are meeting these tools at the precise moment they should be building the very thing the tools are standing in for.

What I am seeing in practice

It looks much the same whether I am in clinic or in a teaching session. An answer arrives within seconds, from OpenEvidence, ChatGPT, Claude, Gemini, or an AI layer sitting over the references we have always used. Ask the question and back comes something fluent, referenced and confident, usually before you have finished forming a view of your own.

I am not against any of this, yet, and I certainly use it myself. But four situations over the past few months have shown me how easily the reasoning drops out of the process, and how difficult that is to notice when the output looks finished.

A medical student named the diagnosis halfway through the first paragraph of a PBL case. They were right. When I asked what else it could reasonably be, and what finding would make them abandon that answer, there was nothing there.

An intern who could recite the systematic approach to a chest X-ray, faultlessly, in order. Then we put the film up and they could not point to the consolidation. They had the framework without the eye, and only one of those can be retrieved.

A registrar who handed me an outpatient letter running to several pages. Immaculate, every guideline referenced. But I read it twice looking for the two sentences that mattered. What do we think is going on with this child, and why. They weren’t there. Our own training was built on long cases: take the full history, then stand up and say what matters, in what order, and what will still matter in five years. The value was never in the volume. It was in the compression. Deciding what to leave out is the clinical judgement.

The clearest one, though, was a doctor working up vancomycin dosing for a complex paediatric patient with renal impairment. The AI evidence tool returned three pathways: the local adult protocol adjusted for age, one adjusted on eGFR, and one built around therapeutic drug monitoring. Laid out cleanly, side by side, like a menu.

To a new player that reads as a choice. Pick one, ideally the one with least friction, and move on. To an experienced one there are not three options. Given this child’s comorbidities and renal function there is one, and it is to call the infectious diseases team who govern vancomycin dosing in paediatrics. The correct answer, was option 4.

The tool wasn’t wrong. Each pathway is defensible in the right patient. What it could not do was tell you which patient was in front of you, or that the correct answer was not on the list. And it was delivered with complete confidence. Recognising the difference is a judgement that sits with the experienced clinician, and if that person is still learning, or was never taught to look, nothing else in the system catches it.

The pattern across all four is the same. The output looks right. The middle is missing. Gather, hypothesise, test, revise, then condense to what matters: that sequence collapses into a single retrieval step at one end, or an undigested pile at the other. We have offloaded arithmetic and drug references for decades without much harm. What has changed is that we are now offloading the reasoning itself, the part that was never a lookup and was always the point of the training.

Which leaves the question:

If we never develop the skill of asking the right question, how will we know whether the answer we have been handed is sound?

What the AI tools can already do

The second paper was from Google Research, on AMIE (Video), their medical AI system extended from text into live video consultation.

They ran 300 consultations across 100 scenarios with trained patient actors, alongside practising GPs completing the same cases. A panel of 20 physicians rated the consultations. The system was rated on par with the doctors on history taking, diagnostic accuracy, management and communication. It rated higher, on average, at eliciting physical signs and guiding patients through examination manoeuvres.

The limitations matter and the authors state them plainly. These were actors in simulation, not real patients with undifferentiated presentations, and only conditions that can be convincingly portrayed. Validation in real practice is the necessary next step.

Even with those caveats, put the two papers side by side and the problem takes shape. What AI can do in a consultation is climbing fast, and it is climbing into territory we assumed was ours for another decade at least. Drawing out a history. Guiding an examination. Making a patient feel heard. Meanwhile the thing our trainees are meant to be building underneath all of that may be thinning out, quietly, in exactly the years we are not looking.

Put next to the Nature Medicine argument, this raises an uncomfortable question about what is in store for us in a few short years. One paper says the analytical foundation may not be forming in the people coming through. The other says the interpersonal work, the part we have always fallen back on as irreducibly human, is not as exclusively ours as we assumed. Squeezed from both directions, what is actually left that is ours?

Quite a lot, I think, but it is narrower and more demanding than the job most of us trained into. Judgement about this particular child, in this family, with this history and these constraints, is ours. Knowing that the correct answer is not on the menu is ours. Carrying the responsibility for the decision is ours. What has changed is their standing. These used to be the things that accumulated quietly around the work of diagnosing and managing. They are increasingly becoming the job description itself. And every one of them rests on exactly the foundation the first paper suggests may not be forming.

Who is responsible for the decision?

The clinician, using a tool they cannot inspect? The vendor, who built the model but has never met the patient? The health system, which procured it and designed the workflow around it?

AHPRA, the Colleges and most of the emerging guidance are converging on the clinician being accountable for the decision to use AI and for what they do with its output. As a starting position that is reasonable, and it is consistent with how we handle every other tool in clinical practice.

But I don’t think it holds indefinitely. These models are developing faster, and are more complex, than any individual clinician can reasonably be expected to keep pace with.

I believe that responsibility will settle wherever competence sits. We cannot reasonably hold a clinician accountable for failing to override an AI system if they were never taught when to override it, or given any basis on which to judge.

Consider how we are trained. Medical school and specialty training are built on understanding what we do to patients. We learn the pharmacology before we prescribe. We learn the mechanism of action, the half-life, the interactions, why a drug works in this child and not that one, and what happens when it doesn’t. The same applies to procedures and devices. Understanding the intervention is the precondition for using it, and we treat that as not negotiable.

AI in healthcare is unique because it is the first instance I can think of where the average clinician, and often the specialist too, does not meaningfully understand how the tool works, and uses it anyway. We are not accustomed to that, and I don’t think we have properly reckoned with what it means.

Which is why I keep landing back on education. Regulation will always be a step behind the technology. Procurement can insist on evidence, but it cannot follow a clinician into the room. What is left is the person making the decision, and whether they were taught enough to make it well.

A clinician who understands what they are using can be held properly accountable, rather than being the last signature on a workflow somebody else designed. A health service full of people like that asks vendors much harder questions before it buys anything. And a system that has already deployed these tools, as ours largely has, at least has a workforce capable of governing them.

Back to the drawing board

There is a line that gets repeated at every digital health conference I attend, usually without attribution. It comes from Curtis Langlotz at Stanford, who said it about his own specialty: AI won’t replace radiologists, but radiologists who use AI will replace radiologists who don’t.

I think that is broadly right, and it is now said about doctors generally. Whether AI replaces clinicians altogether is a separate argument, and one I will park for another day.

So, with that in mind, back to the drawing board.

Which competencies must remain ours, and which can reasonably be delegated? Some dependence on technology is entirely sensible. We need an honest account of where that line falls, rather than a reflexive defence of everything we currently teach.

What does critical thinking look like in the age of AI? Knowing when to accept, question or override an output is a distinct skill, and it does not develop through exposure alone. It is the same critical thinking we have always taught, applied to a source that is fluent, confident and frequently right, which makes it considerably harder to teach. .

How do we teach students and trainees to interrogate the tool itself? Not how to use it, but what to ask of it. What population was it built on. Whether that population resembles the families we see in Western Sydney. What it does when it is uncertain, and whether it tells you. We teach critical appraisal of a trial and expect students to spot a biased sample. We should expect no less of a model.

What do we stop teaching to make room? Curriculum time is finite, and every addition displaces something. This is the conversation nobody wants to have, and avoiding it is how AI content ends up as an optional module in third year that everybody skips.

And who teaches it? Approximately 9% of medical faculty report AI expertise. That is the real constraint on every good idea in this space, and where I expect a good deal of my time will go. It also means modelling it ourselves, showing students our own use out loud, including the times we decide against it.

None of this is an argument for keeping AI out of medical education. Used well, it should prove to be the most powerful teaching technology we have had. But the sequence matters, and at present we are arriving at that sequence by accident rather than design.

If you are working on curriculum, assessment or accreditation in this area, I would welcome the conversation.

Declaration of AI usage: I used AI to help summarise both papers. In my defence, I had already formed a view. That is the entire argument, and I am sticking to it.

End of content

No more pages to load

Log In Register ×