“Human in the Loop” Is Not a Safety Architecture
by Serelora
This article was originally published on Medium.
Read full article on Medium
By Luis Cisneros, CEO Serelora
The doctor in the clinical AI sales pitch has a wonderful job. The software does the work. The doctor exercises judgment. Patients get better care, administrators get their efficiency, and everyone gets home for dinner.
Then you read the instructions. The doctor must verify the work.
All of it, apparently. Carefully enough to catch whatever the software missed, invented, or misunderstood. Quickly enough to preserve the time savings. Consistently enough that the arrangement still works after six months of correct answers and a waiting room full of people who would like to be seen before they die of old age.
This doctor deserves a raise. They seem to be doing two jobs while appearing in the budget as a cost reduction.
“Human in the loop” sounds reassuring because it contains a human. We should probably ask what has been done to that human’s working day.
Healthcare Dive recently reported on the risks of AI scribes producing incomplete or inaccurate notes, alongside warnings that clinicians may scrutinize those notes less as they grow accustomed to using them. The familiar response is to remind doctors to review what they sign. Fair enough. Now show me how that review is supposed to happen. (healthcaredive.com)
An invented sentence at least has the decency to appear on the page. An omission offers no such courtesy. A doctor can read a perfectly coherent note and still fail to notice that a consequential detail has vanished. Finding it may require returning to the original encounter or comparing the note against information elsewhere in the record. That is work. It requires evidence and time, neither of which can be supplied by making the disclaimer more emphatic.
Imagine approving 5,000 correct outputs. At some point, you would reasonably begin to trust the thing. That was presumably the point of buying it. When the 5,001st contains an error, it arrives wearing the same respectable clothes as the previous 5,000.
A product that succeeds by earning trust should have an answer for what happens once it has earned it.
Otherwise, we have built a peculiar arrangement. The machine receives credit for the speed. The doctor remains responsible for the exceptions. The institution can count the minutes saved typing, while the minutes required to establish that the result is safe become a matter of professional conscientiousness. Convenient accounting, provided you never have to perform the review yourself.
The September 15 Nature Medicine paper by Li Zhang and colleagues offers a more interesting approach. The researchers tested a diagnostic AI and evaluated signals of reliability. Agreement across repeated runs of the same case was the strongest signal they examined. At a consistency threshold of 0.90, their proposed approach retained 49.4 percent of cases for autonomous handling, with 98.9 percent diagnostic accuracy in that selected group, and deferred the remainder for review. (pubmed.ncbi.nlm.nih.gov)
These were retrospective simulations. The results do not establish safety in routine care or prove that clinicians would have less work. Errors remained in the selected group. The useful contribution is a method for deciding where review might be needed, which can itself be tested. (pubmed.ncbi.nlm.nih.gov, digitalhealth.tu-dresden.de)
That is a considerably more serious proposition than attaching a doctor to the end of a process and declaring the process supervised.
It also creates obligations. If the system decides what deserves attention, we have to investigate the cases it waves through. Being consistently wrong remains an available option. A reliability score needs evidence behind it, and the consequences of an error should affect how much autonomy a system gets. Permission to draft a note is a poor basis for permission to carry out a treatment decision.
The review itself needs a design. Give the clinician the disputed finding, the supporting source, the contradiction, the missing information. Make it possible to inspect the evidence without an archaeological expedition through the chart. Explain what needs deciding. Give the decision to someone qualified, with enough time to make it.
And account for what happens when that person is unavailable. A message deposited in an unattended inbox is an administrative event. Calling it an escalation does not summon a doctor.
This is where the pleasant language of oversight meets staffing. Who is covering the queue? How long can a case wait? What can the reviewer stop or reverse? Who checks that the decision reached the people and systems that need it?
Those questions cost money to answer. They also make the promised savings harder to calculate. Removing half the cases from a review queue might leave the difficult half, each requiring more attention than the average case ever did. You cannot count heads, count clicks, and assume you have measured thought.
I would want to see the performance of the entire arrangement. Errors caught. Errors missed. Delays introduced. Harm prevented. Cases that should have reached a clinician and never did. Whether the arrangement still works after the novelty wears off and the staff have learned its habits.
A signature is a very easy thing to count. That may help explain its appeal.
What bothers me about “human in the loop” is how readily it turns a person into an abstraction. The doctor becomes an inexhaustible supply of judgment, available whenever the software reaches its limit. Their fatigue, workload, access to evidence, and authority to intervene disappear into four reassuring words.
If your safety argument depends on that person, their working conditions belong in your safety argument.
Put the review time in the budget. Put the evidence in front of the reviewer. Test what happens when the reviewer misses something. Then we can have a useful conversation about safety.
The doctor is already in the building. Tell me what you have built that makes it possible for them to catch the mistake.
RELATED ARTICLES
Explore more insights and perspectives from our team.


