Overview
AI can improve accuracy and efficiency, but it does not eliminate human judgment. It changes where judgment happens. Drawing on lessons from soccer and healthcare AI, this article explores why healthcare leaders should focus not only on model performance, but also on the decisions AI leaves behind and the governance required to manage them.
Key takeaways
- AI does not eliminate human judgment. It relocates it.
- Human oversight becomes more difficult when clinicians are evaluating AI recommendations instead of making the original decision themselves.
- Even successful AI tools depend on human choices about thresholds, workflows, and escalation paths.
- Healthcare organizations should define which clinical decisions AI is allowed to influence before deployment.
- Governance should measure the quality of the decisions AI leaves behind, not just model accuracy.
There is a pattern in football (or as we call it in the United States, soccer) that is worth bringing to your hospital or health system’s next AI governance meeting: Precision does not eliminate judgment; it relocates it.
Watch enough football and you have seen the line. Someone scores, the celebration starts, and then the broadcast cuts to a frozen frame with two bright parallel lines laid across the turf, one at an attacker’s shoulder, one at a defender’s knee. The verdict, made by cameras and computers, comes back in millimeters, deciding that the player was offside. Sixty thousand people groan at something not one of them could see.
The line is a real achievement. It is also a strange object if you look a second longer than the broadcast wants you to. To draw it at all, someone or something must decide the exact instant the ball was played, because offside is judged at that moment and not a second earlier or later. For years, a person picked that frame by hand. Now, 12 cameras track 29 points on every player 50 times a second, and a sensor in the ball reports 500 times a second, so the system can propose the kick point itself. And then an official still checks, by design, that the machine chose the right moment.
People of the internet, you may be surprised to learn the arguments did not end. They just moved. Now we argue about whether a player who barely grazed the ball really played it, and about where a computer decided the edge of a player’s shoulder was.
In football, and in the daily work of delivering care, every relocation of human judgment puts the judgment somewhere smaller, later, and harder to see. Behavioral scientist Michael Hallsworth calls these fractal errors: Technology reproduces them “at smaller and smaller scales, rather than eliminating them.” He maintains that the most urgent test case is not sports, but AI.
Let me grant the obvious objection: Football is entertainment and medicine is not. Nobody mourns the lost spontaneity of a missed diagnosis, and accuracy in our work is close to the whole point, rather than a single value competing with pageantry. Fine. But the mechanism transfers anyway, because it is not about whether accuracy is worth having. It is about where human judgment moves when technology makes part of the process more accurate.
Where human judgment moves
Take a look at your clinical AI tools and ask what they are actually removing. Ambient documentation moves the decision about what matters in a visit from composing a note to auditing one. Predictive alerting shifts the decision to treat to determining whether this alert, on this patient, at this hour, deserves attention. Imaging AI moves the radiologist’s role from forming an impression to deciding whether to disagree with one already on the screen.
Each of those relocated decisions is rarer than the one it replaced, which means less practice. Each is more complicated, because the easy cases were exactly the ones the model absorbed. And each arrives stripped of the reasoning that used to surround it, because you retired that reasoning along with the task.
We have evidence for how that goes, and it is not comforting. In a 2023 study published in Radiology, 27 radiologists read 50 mammograms and provided their Breast Imaging Reporting and Data System assessment assisted by a purported AI system. Twelve of the AI system’s suggestions were deliberately wrong. Among inexperienced mammogram readers, correct assessments fell from 79.7% to 19.8%. Among moderately experienced readers, from 81.3% to 24.8%. Among the very experienced, from 82.3% to 45.5%.
Think about the last pair. Give an experienced breast radiologist a confident wrong answer from an automated decision-making system, and the radiologist’s assessment is correct less than half the time. Two decades of expertise bought only roughly twenty-five percentage points of protection against a machine that was simply mistaken. Whatever we are calling “the human in the loop” in our board decks, this is the actual performance of that human on the actual decision we left them.
They are not naive about it, either. A survey of Norwegian breast screening radiologists found that 47% rated the risk of automation bias as high, while 68% expected AI to improve their own cancer detection. They know the trap exists, and … they expect to be the exception.
Further, the tool does not have to be broken for this to matter. In the Swedish MASAI trial, AI-supported screening found 6.4 cancers per 1,000 women compared with 5.0 for standard double reading, with no rise in false positives and 44% less reading work. It works. Now look at how it works. The AI scores each mammogram, and that score determines whether one or two radiologists read it. Somebody chose where those score thresholds would be. That choice allocates human attention across a national screening program, and it was made before a single radiologist in Sweden opened a single case. A 2025 systematic review states that triage cuts reading volumes by 40% to 90% when thresholds are “conservatively calibrated.” The entire benefit rests on a calibration that no one at the bedside will ever see. That is the kick point, and a person still picks it.
When AI changes human behavior
Football’s other lesson is what happens when you answer residual error with a more precise rule. Handball got clarified, and clarified again, until the sport produced one of the strangest sights in professional athletics: defenders running with their arms pinned behind their backs, giving up a capability so they could not be blamed for having used it.
In healthcare, we call that defensive medicine, and we usually blame the threat of malpractice litigation. But an AI tool can produce the same crouch on its own. Every clinician who has stopped trusting an instinct because the model disagreed, or documented around a recommendation rather than against it, is running with their hands behind their back.
If you’re a healthcare executive, you almost certainly have a dashboard for model performance, but you likely have none for the decision the model left behind. The override rate (an attendance record of how often someone disagreed with the tool) does not count, yet health systems keep treating it as though it does. It tells you nothing about whether they were right, whether they had ninety seconds or nine, or whether they knew something the model could not.
Two governance questions healthcare leaders should answer
I have argued before that a good tool interrupts rarely and earns each interruption, which is what made Penda Health’s AI clinical copilot interesting, and that predictive tools should be validated on your own patients before they go live, which was the governance point in my delivery robot piece last spring. Both of those questions are about whether the tool is any good. Curiously, though, neither of those is the question football answered first. The lesson for healthcare leaders is to spend less time asking whether a tool works and more time governing the decisions it reshapes.
- Write the list of reviewable decisions before deployment. For the first eight years, the laws of football permitted video review in exactly four situations: goal or no goal, penalty or no penalty, straight red card, and mistaken identity. Four. Everything else on the pitch stays with the referee, permanently, by design. That list was written before the first review ever happened. It held for eight seasons. This summer, the list grew, and the arguments about where it should stop grew with it. Your health IT governance committee has an approval pipeline and almost certainly no equivalent document. Which clinical decisions are a machine permitted to touch at all, and which belong to a person because we decided they do? Answer that in writing, in advance, or it gets answered for you one procurement at a time.
Notice the second half of that design. Video review checks every situation automatically and surfaces almost none of them, and then only for a “clear and obvious error.” We built the inverse in healthcare with narrow checking wherever somebody bought a model and unbounded surfacing at whatever threshold shipped in the default configuration.
- Instrument the decision, not the model. If your only measure of human oversight is how often clinicians clicked past the tool, you are counting, not governing. Sample the overrides. Find out what happened to those patients. Ask whether the people doing the overriding had the time and the context the decision required. That is a much more difficult program to stand up than a model-accuracy dashboard, and it is the only one that tells you whether the residual judgment is any good.
Accuracy still requires judgment
The offside line looks objective because it is straight and bright and drawn by a machine, and because the broadcast gives you two seconds to accept it before play resumes. It is genuinely more accurate than the linesman’s eye. It also rests on a choice that nobody in the stadium sees, made by someone whose name will never appear on the screen.
Your model is probably more accurate than the clinician as well. Ask who is picking the frame.
About the author
Craig Joseph, MD, FAAP, FAMIA, is Chief Medical Officer at Nordic and co-author of “Designing for Health: The Human-Centered Approach.” A pediatrician, clinical informaticist, former Epic leader, and former CMIO, he helps healthcare organizations improve patient experience, operations, and technology adoption through human-centered design.