Scale, pause or kill: How health systems are making AI portfolio decisions

Advertisement

The pipeline of AI tools seeking a foothold inside health systems has never been longer or faster-moving. For CIOs and CMIOs, the hardest question is no longer whether to pilot new technology, but how to know when a pilot has earned the right to scale, and when to cut it loose.

At Becker’s 11th Annual Health IT + Revenue Cycle Conference in Chicago, four health system leaders and an AI company executive outlined the frameworks, decision structures and common failure modes shaping how their organizations manage AI portfolios.

Nirav Shah, MD, associate CMIO for AI and innovation at Endeavor Health, a nine-hospital system in Chicago serving roughly 1.3 million patients annually, has more than 100 AI tools in deployment.

“There’s three things that I’ve used probably for the last 10 years, whether it’s predictive analytics or even clinical decision support all the way to LLMs and agents,” Dr. Shah said.

1. The first is model performance. Is the model generating enough signal, and are false positives, hallucinations and bias being adequately managed?

2. The second is technical and workflow integration. Does the live inside the interface clinicians already use? Or will the users have to navigate to a new interface?

3. The third is value creation. Dr. Shah thinks about value creation as either potential or realized. Does the tool solve a problem that matters? Can it actually deliver on that promise?

    “If it doesn’t meet those three lenses, then you either optimize, pause or kill,” he said.

    Dr. Shah offered a candid illustration of what happens when the third criterion is met in one context but fails to transfer to an adjacent one. Endeavor Health had built strong adoption of chart summarization for hospitalists and assumed a similar deployment for case managers would follow easily. It did not.

    “We found that when we deployed it, it didn’t get the value that we thought it was going to get. And we realized that the problem was we had one of the oldest instances of Epic at that time and we weren’t pointing our LLM at the correct notes because we had so much complexity,” he said. “So we paused, waited, cleaned up our entire systems, our note infrastructure and all of that. Now we’ve repositioned it to specific notes and we’re seeing significant value.”

    “Sometimes we all want to go really fast, but sometimes we go fast by going slow and fixing the underlying infrastructure,” Dr. Shah said.

    Doug King, senior vice president and chief digital information officer at Northwestern Medicine — a $12 billion health system anchored by Northwestern Memorial Hospital in Chicago — said two elements are consistently underweighted in pilot design: the clinical champion and workflow standardization across sites.

    “You need a clinical champion. If it’s going into the clinical arena, you just have to partner because the technology is there,” Mr. King said. “But you have to almost be standardized or very close. Because if you’re deploying something at a pilot site and then you want to go health system wide, but they have workflow differences, there’s different ways they document. Those are all going to be very difficult to overcome as well as it’s a maintenance nightmare.”

    The moment a pilot fails to show results is also a governance test. Mr. King said Northwestern Medicine sets success criteria before deployment and holds to them.

    “We agree up front what the value proposition is,” he said. “If it doesn’t [hit those metrics], it needs to be shut down. And we shut it down. We can always revisit it. But for that pilot, there’s 10 more that are waiting in line.”

    Extending timelines without clear evidence can become a trap, Mr. King said. “I think a lot of times what happens is we get into, well, the next version will be better, or just give us two more weeks, that type of thing. The reality is you can do that in perpetuity and never get onto the next thing.”

    When a pilot does need to be terminated, Mr. King said the decision typically involves himself, the CFO, the COO and a senior clinician.

    “You gotta get the right people in the room that are going to be thoughtful, objective and non-emotional,” he said. “Because people get tied to these things because they want it to work.”

    Michelle Gelroth, chief transformation officer at Aspen Valley Health, a 25-bed critical access hospital in Aspen, Colo., said her organization’s smaller scale demands a stricter filter. Volume constraints alone eliminate some technologies before any formal evaluation begins.

    “We don’t want to do AI for the sake of doing AI,” Ms. Gelroth said. “AI for us has to be proven. We can’t be on the bleeding edge. We don’t have the volume to safely test that stuff.”

    Like Northwestern Medicine, Aspen Valley Health routes pilot decisions through a multidisciplinary council that includes clinical and operational representation. But Ms. Gelroth said governance also requires anticipating when a tool will be rendered unnecessary, particularly as Epic continues to expand its native AI capabilities.

    “You have to be able to have a fast review process for that,” she said, adding that the organization has focused on negotiating 30-day no-cause termination clauses in vendor contracts to preserve exit flexibility.

    Edward Sankary, vice president and chief health information officer and chief value officer at UT Health San Antonio, the largest academic medical center in South Texas, said the strongest pilot decisions are anchored in a clearly defined problem, not a technology opportunity.

    “Ambient technology was physician burnout,” Dr. Sankary said. “So it wasn’t just we got a really cool technology, but it was solving a clear problem and it integrated with the workflow.”

    Dr. Sankary also flagged a failure mode specific to patient-facing AI: the organization piloted AI-generated responses to patient medical advice requests and found that it added to physician time instead of shrinking it. Physicians had to read both the patient’s question and the AI-generated response before composing their own reply. The tool was paused.

    Dhruv Chopra, CEO and founder of CIVVI and RAD Pod, offered a different frame for the kill decision: whether a tool moves the organization closer to full autonomy or only partway there.

    “We don’t want to create something that’s just going to give us a copilot and it’s going to create a dashboard and it’s going to involve people looking for signals,” Mr. Chopra said. “We look at it and say, what is going to get us closest to autonomous?”

    Mr. Chopra’s organization, which he said has replaced roughly 80% of its workforce with AI agents running 24 hours a day across scheduling, radiology interpretation, revenue cycle coding and other functions, evaluates potential investments primarily on the degree to which they eliminate process friction end to end rather than improving a single step.

    For health systems still building the governance muscle to apply any of these frameworks consistently, Dr. Shah offered a closing note on what makes scale possible in the first place.

    “AI is going to scale at the speed of trust,” he said. “The trust is really: are we solving the key problems? Is the model performing well? Is the workflow seamless and are we creating that value? You have all of that, then you can scale.”

    Advertisement

    Next Up in Innovation

    Advertisement