Keeping its vast network of 145 million wireless users and 16 million broadband customers connected is paramount for AT&T. A brief 20-minute service disruption can lead to significant financial repercussions for the company, as it frequently offers bill credits to customers during service lapses.
While most outages are localized and generally fly under the radar, there are occasions when they impact millions of users. For instance, in February 2024, a significant wireless outage affected over 92 million voice calls and interrupted more than 25,000 911 calls, as reported by the US Federal Communications Commission. Other leading telecommunications companies like Verizon and T-Mobile have similarly faced substantial service interruptions in recent years.
In an effort to enhance outage responses and customer communication, AT&T embarked on a mission to centralize its network interruption detection system. "We need to be aware of any network issues immediately and resolve them, ideally before customers notice," stated Andy Markus, AT&T's chief data and AI officer, in an interview with Business Insider.
The initiative began in 2017 with the development of an end-to-end incident management system (EEIM), leveraging technology from various providers, including MongoDB and Snowflake. Initially focused on conventional machine learning techniques, AT&T later incorporated generative and agentic AI features, broadening the system's capabilities, according to Markus. This tool has been effectively utilized to manage service disruptions impacting both individual customers and small enterprises.
Currently, the EEIM employs a combination of AI-driven and machine-learning techniques to determine the underlying causes of disruptions, prioritize issues, suggest remote fixes, and alert customers proactively when outages occur, Markus explained.
Telecom outages can arise from numerous factors such as severe weather, technical failures, cyberattacks, and physical infrastructure damage. When customers experience connectivity issues with their mobile or landline phones, or internet services, they often inundate AT&T’s customer service departments with requests for assistance.
In early 2017, AT&T recognized the need for a more agile response to outages. The company created a cross-functional team to garner input from employees across various departments, including IT, data analytics, and AI, with the intent of aligning the new EEIM with existing network technologies while also integrating innovative AI tools into their operations.
Frontline technicians, responsible for network maintenance and repair, were also consulted, as were the network teams that monitor traffic and swiftly restore services during interruptions, Markus noted.
The EEIM enabled AT&T to reorganize an extensive dataset of 10 petabytes, equivalent to 5,000 billion pages of text. This data encompassed network logs, alarm signals, incident dispatch records, and details about outages. Markus emphasized that consolidating and restructuring this data was critical for identifying issues, predicting their occurrence, directing responses, and supporting new features.
MongoDB, established in 2007 as a document database provider, played a vital role during this phase. Its platform is designed for scalability, allowing AT&T to expand storage capacity without the need for a complete system overhaul, referred to as "sharding." This approach facilitates data distribution across multiple systems to handle extensive datasets, explained Ben Cefalo, MongoDB's chief product officer.
Additionally, the EEIM relies on Microsoft Azure for cloud computing, combined with data analytics from Databricks and incident reporting through Snowflake, Markus stated.
According to an AT&T representative, the EEIM was first implemented for broadband fiber services in early 2018 and expanded the following year to include digital subscriber lines (DSL).
In June 2018, AT&T introduced a dedicated application called Atlas, utilized by field technicians. It integrates machine learning and AI models to diagnose outages and recommend remediation plans.
AT&T had already been using these traditional AI methods prior to Markus's arrival in July 2020, but his leadership accelerated the adoption of AI across the organization. Currently, around 100,000 employees utilize AT&T's generative AI tools, processing more than 27 billion tokens daily and refining smaller language models to enhance cost management.
By early 2021, the company also launched a proactive customer notification system through the EEIM. "This really alleviates customer stress and frustration," Markus remarked.
With the introduction of generative AI features in the EEIM during the first quarter of 2022, the system began to analyze historical data to identify the root causes of outages. In early 2025, AI agents were incorporated into the platform, enabling them to gather outage information from customers, resolve issues, and provide case details to technicians when on-site intervention is needed.
Markus noted that AT&T has developed over 30 AI models to forecast potential problems like configuration errors, weather impacts, and system failures, allowing the system to act preemptively. "This capability helps us pinpoint active issues or anticipate future challenges," he explained.
Overall, the AI-enhanced EEIM has allowed AT&T to avoid 3.1 million unnecessary field dispatches and reduce customer downtime by more than 12 million hours in the past year, Markus concluded.



