In my observation of healthcare AI pilots, nearly all that meet their success criteria end up facing significant challenges once they transition to a full production environment, often within a few months. These obstacles typically arise from factors unrelated to the AI model itself. Unfortunately, we continue to assume that the results from pilot projects will accurately reflect performance in the hospital's operational context, which is far from the truth.
The success of these pilot programs is often attributed to the acceptance of design and architectural shortcuts that cannot be replicated in real-world production settings. For instance, data may be pulled in a one-time extraction for the vendor, or a custom integration might be quickly assembled over a weekend. Nightly refresh jobs may be initiated simply because negotiating API quotas with EHR vendors takes too much time. Consequently, the clinical champion showcases a polished demonstration, the budget owner yields a seemingly feasible ROI, and the production rollout receives the green light. Fast forward eighteen months, and with several AI tools operating simultaneously, issues arise: integration teams become perpetually on call, unexpected API quota issues emerge, and clinicians find discrepancies between how different AI tools interpret the same patient data. These challenges do not stem from the AI technology itself.
By 2025, healthcare AI expenditures rose to around $1.5 billion, with EHR vendors releasing numerous AI functionalities. However, the rapid adoption of these tools has outpaced the necessary architectural foundation. Healthcare organizations are acquiring AI in a manner reminiscent of the past EHR procurement methods, making individual decisions independently without a cohesive data governance framework. The accumulation of integration debt from this approach mirrors that of the EHR implementation phase, yet the current situation is compounded because each AI application interacts with the same underlying data upon which subsequent tools depend. As a result, the associated technical debt is increasing at an unexpected rate, making future cleanup more complex.
Three recurring issues complicate these scenarios, each of which is cheaper to address proactively than to remedy later.
First is the data-on-demand fallacy. AI vendors often assert their tools can seamlessly integrate with existing data, a claim that holds true during the pilot phase. However, when health systems begin deploying multiple AI tools, each vendor typically establishes their own data pipeline that interacts with the EHR, complete with unique extraction procedures and field mappings. This proliferation means that integration teams will ultimately be burdened with maintaining these individual pipelines indefinitely. In turn, EHR API call limits fluctuate unpredictably, and discrepancies emerge in the data each tool accesses. Consequently, one tool might indicate a patient has been on a medication for an extended period, while another reports they just started it. The issue stems not from individual inaccuracies, but from each tool accessing different snapshots of the same patient data. The solution lies in creating a well-governed data layer—preferably FHIR-native—situated between the EHR and the AI tools, ensuring all applications draw from a consolidated data source. Unfortunately, most health systems have not yet established this essential infrastructure, and delays will necessitate considerable migration efforts when the time comes.
Next is the governance lag. Awareness of AI governance has surged in healthcare recently, yet tangible improvements remain minimal. Many health systems can identify the AI applications they use, but few can address critical questions. Who is liable when an AI-generated transcription contains a medication error? What approval processes are in place for the specific model versions employed in various tools? How do organizations handle audits from regulatory bodies for AI-generated outputs? Effective governance must extend beyond individual applications to the shared data and integration layer that supports all systems. Key functions such as model versioning, data lineage tracking, and audit capabilities must exist at the platform level rather than within each tool. When changes to governance are retrofitted post-implementation, issues arise with model updates, raising the risk of compliance problems.
Another concern is the so-called agentic ceiling. For many health systems, procuring generative AI solutions has become routine—select a vendor, conduct a pilot, and proceed with implementation. However, agentic AI, which actively engages with patient records, presents a different challenge. These systems require real-time integration, transactional reliability, and robust authorization protocols that can maintain performance under demand while providing comprehensive audit trails for every action taken. Initiatives like CMS’s WISeR program are directly integrating AI-facilitated prior authorization into traditional Medicare workflows, prompting EHR vendors to release agentic features across various operational domains. Many health systems already face strain from generative AI applications; when agentic capabilities become standard through vendors without the necessary infrastructure, the ensuing challenges will manifest quickly.
This scenario is not speculative. By 2027, we can expect pilots that performed well in 2025 to experience significant slowdowns in their production environments. Integration teams could become overwhelmed managing their myriad point-to-point connections, leading to incidents that prompt elevated scrutiny from leadership. Organizations in such predicaments often find the path back to stability requires extensive time and resources, usually exceeding the initial investment required for a robust architecture. Similar situations unfolded during EHR implementations in the 2010s, with significant cleanup work being largely avoidable.
The solution is not glamorous, but it is essential. Health systems must invest in their data and integration layer before approving new AI procurements. Establishing a clean, FHIR-native data layer to connect the EHR with AI tools is crucial, as well as mandating that all new vendors utilize this centralized data source instead of creating independent pipelines. Governance functions—including model versioning, lineage tracking, and auditing—should reside at the platform level rather than within each application. Additionally, it is vital to treat agentic functions differently from generative ones in terms of architecture.
The upcoming eighteen months will distinguish between two types of health systems. Those that prioritize a strong data and integration foundation now will find that their subsequent AI investments build effectively on this framework. Conversely, organizations that keep layering new AI tools onto a fragile architecture will face a daunting future of overdue foundational work, which should have been established by 2026. Ultimately, the complexity of AI implementation is not rooted in the technology itself.



