What should your hospital sound like when a nurse is worried?

By:

Craig Joseph, MD
ambient AI
Overview

Ambient AI and automated documentation may eliminate documentation signals that reflect clinical judgment. Because those signals can help predict patient deterioration and inform AI models, health systems should understand what may be lost before automating clinical workflows.

  • Clinical documentation patterns often contain valuable signals of clinician judgment, not just patient data.
  • Research shows nursing behavior and documentation activity can predict patient deterioration and improve outcomes.
  • Many predictive AI models rely partly on clinician behavior, creating risk when workflows change.
  • Ambient AI and automated documentation may eliminate documentation “breadcrumbs” that currently serve as clinical signals.
  • Health systems should identify which signals automation removes and intentionally preserve those that support patient care.

Nobody ever designed cars to be loud. The noise was a byproduct of burning gasoline, and engineers spent decades trying to muffle it. Then hybrids and electric vehicles made cars quiet, and we discovered that engine noise had been serving as a safety signal all along. Blind pedestrians use the sound of individual vehicles to determine where a car is, how fast it’s moving, and which way it’s headed. In 2016, federal regulators required quiet vehicles to emit an artificial sound below about 19 miles per hour, because removing that signal created new risks.

Healthcare may be approaching a similar moment. As ambient AI, automated documentation, and continuous data capture reduce the manual work clinicians do in the EHR, they may also eliminate subtle behavioral “breadcrumbs” that reveal clinical judgment. Those breadcrumbs are often invisible to humans, but increasingly visible to algorithms.

In a recent Becker’s interview about Mayo Clinic’s ambient documentation tool for nurses, Mayo’s Vice Chair of Nursing Practice Transformation, Cheristi Cognetta-Rieke, DNP, RN, said about documentation latency: “I don’t actually care about latency anymore because the conversation is captured.”

On the narrow point, she’s right. Latency mattered because nurses had to hold assessments in their heads for hours. The next question, though, is: What else goes quiet when the chart starts writing itself?

Why documentation timestamps may capture clinical judgment  

In January, I wrote that healthcare had been built around “context-free abstractions (think codes, notes, and timestamps).” I was half wrong, and the timestamp is the part I got wrong. #oops

In 2018, a Harvard team analyzed a year of records covering 669,452 patients at Brigham and Women’s Hospital and Massachusetts General. They compared lab test results against how they came to be ordered: hour of the day, day of the week, interval since the last one. Among the 174 tests where either kind of information predicted anything about three-year survival, the ordering process beat the result in 118. The most predictive variable across all tests they modeled was not a lab value; it was the time elapsed since the last draw.

Think about that for a moment. A result is one indicator of health, but “the fact that it was ordered at 4 a.m. captures the physician’s experience, intuition, and assessment of the patient’s main complaint, baseline status, and physical exam, which are usually not explicitly coded elsewhere.” That’s not noise to be scrubbed out of the data. Indeed, it’s judgment confused as a timestamp.

Nursing documentation behaves the same way, and the effect has been visible in the data since at least a 2013 study in which patients who died had more unrequired vital signs recorded and comments added in their last 48 hours than patients who survived. Informaticists call this “informative presence”: whether data exist, and when, is itself information. The rest of us might call it clinical judgment leaving footprints.

What the CONCERN study reveals about nursing judgment

Researchers at Columbia and Mass General Brigham built an early warning system called COmmunicating Narrative Concerns Entered by Registered Nurses (CONCERN) based solely on those footprints: the frequency and pattern of nursing surveillance documentation. In a pragmatic trial published last year in Nature Medicine, 74 units across four hospitals were randomized, covering 60,893 encounters. Patients whose teams could see CONCERN scores had a 35.6% lower instantaneous risk of death and an 11.2% shorter length of stay. In the authors’ words, they were “a third less likely to die and a quarter more likely to be transferred to intensive care.” Hospice and do not resuscitate/do not intubate patients were excluded, so this isn’t a finding about the end of life. It’s about the patients we expect to send home.

It’s worth being precise about why it worked. CONCERN didn’t outthink the nurses. As the trial team wrote, nurses document additional data to flag a concerning change, “but the patterns of these additional data are not explicitly evident to other members of the care team.” The system put one risk score in front of every nurse and physician on the team. It took one nurse’s worry, expressed as unrequired vital signs at 3 a.m., and made it legible to everyone else.

Dutch researchers arrived at the same place from another direction. In a prospective study of 3,522 surgical patients, a vital-signs early warning score predicted unplanned ICU transfer or unexpected death with an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.86. Adding nine structured indicators of nurses’ worry raised it to 0.91(for those of you who didn’t major in math, an AUROC of 0.5 means the prediction is accurate only half the time. An AUROC of 1 means the prediction is always correct). One of those indicators, “change in behavior and/or doesn’t look good and/or look in the eyes,” carried an odds ratio of 14.6, essentially tied with a change in breathing at 15.2. A nurse thinking a patient looks wrong was worth about as much as the respiratory exam with respect to predicting patient decline.

Then there’s the finding nobody designed. Among patients who went on to deteriorate, 29% had incomplete vital sign sets; among those who didn’t, 76% did. Respiratory rate was missing in 22.5% of the first group and 70.3% of the second. Nurses completed the whole set on the patients who worried them. Chart completeness was itself a signal, on an ordinary surgical ward, with no algorithm in sight.

That changes what extra charting is. We file it under documentation burden, and sometimes that’s exactly what it is. But the unrequired set of vitals is also a nurse telling colleagues, “I’m not comfortable with this patient,” in the only language the chart speaks. CONCERN works because a system finally started listening.

How predictive AI models rely on clinician behavior

And now for the less comfortable corollary. If clinician behavior carries that much signal, some of the models health systems have purchased are running on it. A 2021 study trained models on 42.9 million hospitalizations using an administrative database with no vital signs, no lab results, and no notes, only what was ordered, done, and billed on the first day. They predicted in-hospital death with an AUROC of 0.89, against 0.95 for models built on the full record. Models at that level, the authors conclude, “may derive their signal by looking over clinician’s shoulders,” using clinical behavior “as the expression of preexisting intuition and suspicion.”

That isn’t a scandal. Amplifying clinical judgment is a legitimate thing for software to do, and CONCERN shows it can save lives. But it means these tools depend on clinicians continuing to behave the way they did when the model was trained. Change the behavior and the model drifts. A 2021 New England Journal of Medicine letter named this for clinicians as dataset shift, “a mismatch between the data set with which it was developed and the data on which it is deployed.” Michigan Medicine (the University of Michigan’s health system) had to switch off its Epic sepsis model in April 2020 because the pandemic changed who was arriving with a fever. The next line might worry you more: That was an extreme case, and “many causes of dataset shift are more subtle.” Changes in clinician behavior are on their list.

Now look at the roadmap, including ambient nursing documentation, continuous vital sign capture, and AI-drafted notes. Each is defensible on its own. Each also changes the behavior that generates the breadcrumbs. When the chart fills itself in on schedule, the extra set of vitals at 3 a.m. never registers as a choice anyone made, and the blank respiratory rate that used to mean “this one is fine” gets completed for everybody. Nobody has published on whether ambient documentation weakens signals like CONCERN’s. That absence is the point.

How health systems can preserve clinical signals during AI adoption

Here are two concepts that we should think about to stay on track.

  1. Run a breadcrumb inventory before any automation goes live. G.K. Chesterton advised against tearing down a fence until you know why someone put it up. The questions are simple: What clinician behavior does this tool eliminate, what was that behavior telling us, and which of our models, alerts, and dashboards might depend on it? Ask vendors which features come from what clinicians do rather than from the patient’s physiology. If they can’t answer, you’ve learned something important about the model.
  2. When a signal will be lost, rebuild it on purpose, the way carmakers added the artificial hum. If ambient capture makes nurses’ worry invisible in documentation patterns, give nurses an explicit, low-friction way to register it inside the new workflow, and route it to the care team rather than to a management dashboard. The Dutch indicators are a reasonable starting point, and “doesn’t look good” has better numbers behind it than most of what we build alerts on. That routing decision matters more than it seems. The same breadcrumbs that let a team rescue a patient can be used to count keystrokes. Clinicians will know which one you built.

What healthcare should preserve as documentation becomes automated

Regulators didn’t make electric cars loud again. They decided what a car should sound like and required it. Health systems are about to make the same decision about clinical documentation, whether or not they realize it. Ambient AI will make charting quieter, and it probably should. The question is what your hospital should sound like when a nurse is worried, and whether anyone will still be able to hear it.

About the author

Craig Joseph, MD, FAAP, FAMIA, is Chief Medical Officer at Nordic and co-author of “Designing for Health: The Human-Centered Approach.” A pediatrician, clinical informaticist, former Epic leader, and former CMIO, he helps healthcare organizations improve patient experience, operations, and technology adoption through human-centered design.

Stay up to date on how healthcare’s changing and how we’re helping organizations change with it.

Join us for a night of networking

Join Nordic for an after‑hours networking happy hour at HIMSS. Connect with your chapter’s industry experts over great drinks and insightful conversation. This complimentary event is open to members of all HIMSS chapters.