Health systems are more committed than ever to AI deployment, with most hospitals across the country now using third-party tools for various clinical and administrative workflows. But new research shows that this enthusiasm has outpaced the governance and infrastructure needed to effectively manage this AI surge.
The report — released this month by UPMC’s Center for Connected Medicine and KLAS Research — surveyed health system leaders and found that less than half of hospitals have a dedicated environment for testing AI tools before they touch patient care.
While most organizations evaluate AI solutions before they’re rolled out, those validation methods vary widely, from formal vendor testing to loosely structured pilot programs.
The report also found that 63% of health systems describe their AI strategy as still developing or ad hoc. This gap in testing infrastructure is rooted in time and resource constraints that many hospitals face, according to Ken Howard, vice president of technology services engineering at UPMC Enterprises.
When a hospital identifies a challenge it believes AI can solve, it often doesn’t have the time, capital or talent needed to build a structured testing environment first — so the organization defaults to a standard IT implementation process instead, he explained.
Without a reliable and repeatable way to validate tools, health systems often commit to a typical implementation timeline of six months or more, only to discover the AI solution doesn’t deliver the value they expected, Howard added.
“Without having a dedicated or consistent test environment strategy, they’re going to go down that path just to learn that all that work potentially wasn’t justified,” he remarked.
UPMC has tried to help solve this problem with Ahavi, its real-world data platform that lets the hospitals validate third-party AI tools against de-identified patient data before they’re formally deployed. That kind of pre-deployment testing is only part of the equation, though, said Rob Bart, UPMC’s chief medical information officer.
He noted that governance must extend well beyond a solution’s initial rollout — adding that UPMC has had a formal AI governance structure in place for more than two years, which involves the continuous monitoring of tools after they go live.
This ongoing monitoring is especially important for clinical algorithms, like those predicting hospital length of stay or readmission risk, Bart pointed out.
“We monitor that on regular intervals post-implementation to make sure that the guidance that it is intended to provide is still accurate and reflective of the original [validation],” he declared.
UPMC also tests vendor algorithms against its own patient population, rather than relying solely on a vendor’s testing data, Bart added. This helps catch issues, including bias and model drift, that a more generic data set might miss.
This level of scrutiny is only becoming more important as the pace of AI adoption continues to accelerate, said Kate Eisenberg, senior medical director of DynaMed, an AI-powered tool for clinical decision support.
Eisenberg pointed to a survey from the American Medical Association that showed clinicians’ use of AI tools nearly doubled from 2023 to 2026 — a pace of change that makes it that much harder for evaluation standards to keep up.
As providers increase their scrutiny, she thinks they nee to include a focus on equity.
“We’ve always had in our web interface the opportunity for users to flag if there was an equity concern or not,” she said, noting that clinical teams should be specifically trained to evaluate AI responses for bias.
Having built-in evaluation methods is vital, Eisenberg noted, because the industry still lacks a consistent approach for assessing these tools.
Until standards catch up, health systems and vendors are largely left to govern AI on their own.
Photo: Panya Mingthaisong, Getty Images