Artificial intelligence (AI) is significantly transforming the field of medicine. Its applications now extend from assisting in diagnostic decisions and supporting patient interactions to automating clinical documentation processes. This shift highlights the rapid advancement of machine-learning models and their incorporation into everyday clinical practices. One area in healthcare that holds great potential for AI integration is interpreter services. However, there’s still uncertainty about whether academic research has adequately addressed the new patient-facing AI applications as they emerge.
To explore this, we utilized four AI-based knowledge platforms—ChatGPT, Perplexity, Gemini, and OpenEvidence—to conduct structured queries aimed at estimating the volume of publications over the last five years related to AI in healthcare. We refined our search to focus on interpreter services and also examined literature that included patient perspectives, experiences, or reported outcomes. Research that met all three criteria was categorized as “all criteria.” Our findings indicate a substantial growth in publications concerning AI in healthcare, rising from approximately 11,500 in 2019 to over 28,000 in 2024. Notably, research that specifically investigates how AI affects services critical to patient experience, like interpreter services for non-English language preference patients, is still limited.
The findings underscore a vital need for further research and policymakers' attention. If AI communication tools are deployed without comprehensive evaluation in diverse linguistic and cultural settings, they may inadvertently exacerbate existing disparities in healthcare.
Recent studies have identified risks associated with communication technologies, including significant inaccuracies in automated speech recognition (ASR) for various dialects and accents, as well as inconsistent performance from neural machine translation (NMT) systems across different languages and clinical situations. A systematic review found considerable variability in translation accuracy for NMT technologies in healthcare, emphasizing the importance of thorough evaluations prior to their use in clinical environments. Given the current landscape, there is no established method to guarantee the accuracy of AI interpretations, making the development of clear evaluation metrics essential as new technologies emerge.
AI-driven language access tools, such as ASR, real-time interpretation systems, and video avatars, are gaining traction as alternatives to traditional interpretation methods that typically rely on live human interpreters via phone or video calls. Understanding the roles of these technologies necessitates distinguishing between translation—mainly focused on written text—and interpretation, which involves real-time spoken language exchange. While some AI systems can provide real-time communication, they often depend more on machine translation than on the nuanced human interpretation required in many medical scenarios.
Surveys conducted among clinicians reveal rising expectations for diagnostic AI tools alongside a need for smooth integration into existing workflows. Although these assessments don't directly evaluate AI in interpretation, they highlight the swift adoption of technology and the imperative for rigorous evaluation frameworks before deploying such tools in critical communication contexts. While AI interpretation tools present opportunities for cost savings and improved access in resource-limited clinical environments—especially for less commonly spoken languages—their effectiveness in complex clinical situations remains underexplored. Evidence indicates that AI tools are most proficient in straightforward, low-risk interactions. However, real-world clinical discussions often involve intricate conversations about diagnoses or treatment options, where capturing tone and emotional nuances is paramount.
Currently, AI language tools may be best suited for specific use cases, such as translating low-risk written patient materials or providing temporary support until a live interpreter is available. Even in these scenarios, careful evaluation for accuracy and patient acceptability is essential.
Despite the sharp rise in AI-related research in healthcare, studies focusing on patient perspectives are quite sparse. Existing studies indicate that while patients acknowledge the potential of AI, there remain significant concerns about its safety and effectiveness. Our recent queries revealed that fewer than 0.4 percent of AI healthcare publications address patient perspectives, with virtually no research on interpreter services including this vital aspect. Outputs from AI models showed limited entries related to AI and interpreter services, highlighting a systemic oversight in aligning the rapid technological advancements with the necessary evaluations of their impact on patients with non-English language preferences.
Language is a crucial factor in effective healthcare communication and equity. Patients with limited English proficiency often face higher risks of misdiagnosis and misunderstandings of treatment plans. Therefore, accessible interpreter services are not only integral to providing high-quality care but are also federally mandated in many cases.
Evidence clearly shows that certified medical interpreters enhance communication accuracy and reduce the likelihood of clinical errors. As the demand for efficient healthcare rises, relying on unproven AI interpretation tools instead of trained human interpreters could lead to miscommunications during critical interactions.
To ensure that AI interpreter services are effectively integrated into clinical practice, a patient-focused research agenda is vital. This should include rigorous assessments of accuracy and safety, comparative studies against certified interpreters, and evaluations of how patients perceive and trust these AI systems. Additionally, usability assessments must determine how well these tools fit into clinical workflows and their acceptance among both patients and providers.
Continuous error monitoring systems are necessary to detect clinically significant mistakes in real time, ensuring that any issues can be immediately addressed. There is also a need for community involvement in the creation and oversight of these technologies to ensure that they meet the needs of diverse populations effectively.
In conclusion, the success of AI in healthcare, particularly in interpreter services, hinges on not only patient-centered research but also on institutional frameworks that manage potential algorithmic risks. The role of physician-informaticists and health-system leaders will be crucial as these technologies progress and shape patient experiences in healthcare settings.



