In a recent episode of the medical drama "The Pitt," AI-driven speech-to-text technology took center stage, highlighting its potential in the healthcare sector. The character Dr. Baran Al-Hashimi proclaimed, "Studies show that you can spend 80% less time charting," while demonstrating a new AI tool to her fellow doctors. However, her colleague voiced skepticism when the AI misidentified a patient's medication, raising a pertinent question: Are the benefits of AI tools in medicine worth the inherent risks?
Nelly Elsayed, an associate professor at the University of Cincinnati, sought to address this concern and propose strategies for developing safe and effective AI systems in medical documentation. Her recently published paper, titled "Socio-technical risks of clinical speech-to-text systems: Transparency, privacy, and reliability challenges in AI-driven documentation" in the International Journal of Medical Informatics, delves into existing research, ethical considerations, and regulatory frameworks to assess the rapid advancement of AI against the need for oversight.
While the quality of AI tools continues to improve, Elsayed emphasized that additional factors must be addressed to enhance their efficiency and transparency. She identified five primary risks associated with clinical speech-to-text applications:
1. Inconsistent practices in disclosure and consent 2. Reduced effectiveness when processing accented and disordered speech 3. Background noise in clinical environments compromising AI accuracy 4. Insufficient human review of AI-generated text, leading to unchecked errors 5. Ambiguity regarding accountability for mistakes: Is the responsibility with the software or the clinician?
According to Elsayed, incorporating human oversight in the review process before finalizing data could significantly mitigate these issues. "We need to have a human in the loop to verify that the text accurately reflects what was spoken," she stated, underscoring the importance of thorough checks rather than limiting reviews to just the initial statements.
She also pointed out that various real-world elements can hinder an AI's ability to record accurately. Most AI systems are trained in idealized environments, devoid of the typical distractions found in busy clinical settings, such as machines beeping and conversations among healthcare providers. For AI to prove reliable in practice, it must be trained to understand diverse accents and speech variations.
Additionally, ensuring that clinicians receive adequate training on the software before implementation is crucial. "The organization developing the system must provide clear guidelines to doctors on effective use, potential pitfalls, and what to monitor," Elsayed advised.
In summary, while AI speech-to-text technologies hold great promise for streamlining medical documentation, careful attention to transparency, human oversight, and real-world application is essential for their successful integration into healthcare practices.



