Every healthcare organization has an AI story right now. Fewer have an honest one.
MRO recently convened a focus group of healthcare IT and informatics leaders (CIOs, data executives, and clinical informatics leads from health systems, rehab networks, and home health organizations) through CHIME, to talk candidly about how clinical data is really being used, managed, and activated across their organizations. One of the more useful threads that emerged wasn’t about what AI could do. It was about the gap between what AI is being sold on and what it’s actually delivering.
Healthcareโs AI Sales Pitch vs. The Reality
Ambient AI documentation tools have become nearly ubiquitous across health systems in the past two years, and the pitch that often accompanies them is compelling: automate the note-taking, free up clinician time, see more patients per day.
One participant, a CIO overseeing an active ambient AI rollout across physical therapy, occupational therapy, and speech therapy teams, pushed back on that framing directly:
“There’s this idea that we get our documentation faster so we can see more patients. And I’m not sure that that’s the value of the ambient tool as much as decreasing cognitive load, making it easier to complete the documentation and move forward. So the hype gets sold as, ‘they’re going to have this much extra time, they can see three more people in a day.’ That’s the hype part.”
This distinction matters more than it might first appear. “See more patients” is a productivity claim that’s easy to measure and easy to oversell. When it doesn’t materialize, it undermines trust in the tool itself. “Reduced cognitive load” is a more modest, more honest claim, and it’s also the one clinicians are more likely to actually feel and believe.
For healthcare leaders evaluating or defending AI investments, this is a useful reframe for setting expectations internally: measure and communicate the value you can actually substantiate, not the value that’s easiest to put in a slide deck.
Where AI Chart Summarization Actually Adds Value
Not everything in the AI conversation was skepticism. One participant described a chart summarization feature, using large language models to pull relevant clinical information from notes and discrete data, that had reduced chart-search time for physicians and advanced practice providers by 13% since rollout.
But the more interesting finding wasn’t the time savings. It was what happened when the tool got something “wrong”: “Sometimes they said, ‘oh, this data point’s wrong,’ but then when they actually went to the reference, it was pulling from a note where they’d copy-forwarded information. It was the underlying data that was wrong in their note from the day before. So people are understanding more the value of having correct information in their note, because that’s what the tools are using to create the summaries.”
In other words, the AI wasn’t wrong. It was accurately surfacing an error that had existed in the documentation all along, one that a human reading the same note might have skimmed past. The tool’s real value here wasn’t just speed; it was making an existing data quality problem visible enough that clinicians started correcting their own behavior.
That’s a meaningfully different kind of AI value than “faster,” and arguably a more durable one. It improves the underlying data over time, not just the interface sitting on top of it.
Why AI Is Only as Reliable as the Data Behind It
One participant offered a useful counterpoint. Their organization is using AI to pull diagnosis codes forward from acute care data, a genuinely exciting use case for care continuity. The catch: “Currently the AI is pulling over codes that have already been ruled out, and yet we don’t always know that when it comes in, so we have to look closer.”
This is a small, specific example of a larger pattern that came up repeatedly in the conversation: AI tools are frequently only as good as the underlying data pipeline and documentation discipline feeding them. When that foundation has gaps (stale codes, copy-forward errors, inconsistent data definitions across departments), AI doesn’t fix the gap. It can just as easily amplify it faster and with more apparent confidence than a human would.
What AI Trust Really Depends On
Recruitment is shifting toward real-time, data-driven workflows that start with clinical data.
Instead of relying on retrospective searches, modern approaches integrate with EHR systems to identify potential participants as they receive care. This allows teams Across every AI example discussed, such as ambient documentation, chart summarization, and diagnosis code automation, the same theme kept surfacing: trust in AI output is inseparable from trust in the data underneath it. The organizations getting real value out of AI weren’t the ones with the flashiest tools. They were the ones that had already done the less glamorous work:
- Validating data quality
- Piloting carefully with engaged clinicians
- Being honestโinternally and with vendorsโabout what a tool actually delivers versus what it was pitched to deliver
That’s not a particularly exciting message to put on a conference slide. But based on what we heard from the leaders actually running these programs, it’s the more accurate one, and probably the more useful one for any organization deciding what to invest in next.
It’s also a lesson we’re learning firsthand. Dave Costenaro joined MRO as our lead principal AI architect, and he’s spent his time here doing exactly the kind of foundational work described above: getting into the details of how AI is actually being implemented across our own systems, not just how it’s being marketed. We’ll be publishing more of what we learn along the way, including the parts that don’t go as smoothly as a demo would suggest.
This post draws on insights from a CHIME-facilitated focus group of healthcare IT and informatics leaders convened by MRO. Participant identities have been kept anonymous by mutual agreement.
Frequently Asked Questions
Does ambient AI documentation help doctors see more patients?
Not directly. CIOs running these rollouts report the real benefit is reduced cognitive load and faster note completionโnot more patients per day. Framing it as a productivity gain oversells the tool.
Why is my AI chart summary showing wrong information?
It’s often not wrongโit’s surfacing an error already in the note, like copy-forwarded data from a prior visit. The fix is correcting the source documentation, not the AI.
Is AI reliable for pulling forward diagnosis codes?
Only as reliable as the data behind it. AI can carry forward codes that have already been ruled out if the source system hasn’t caught upโhuman review is still required.
What makes healthcare AI actually work vs. just hype?
Clean underlying data and honest expectation-setting, not the sophistication of the tool. Organizations that validate data quality first see more durable results than those chasing the newest feature.