Translating a single page from 12th-century Latin can require years of expertise, and the task of transcribing a manuscript that spans hundreds of pages can occupy a researcher’s entire career. However, a collaboration between Inria (the French Institute for Research in Computer Science and Automation) and experts in philology has led to the creation of an automated recognition system that successfully processed 32,763 medieval manuscripts in just four months.
This initiative, known as CoMMA—Corpus of Multilingual Medieval Archives—is now freely available online, allowing users to access digitized pages alongside their transcriptions.
Recognizing handwritten text poses significant challenges for machines, particularly with historical documents where the appearance of letters and words has changed dramatically over time. According to Thibault Clérice from Inria's ALMAnaCH team, deciphering hastily written notes or administrative records is notably more complex than interpreting elegantly crafted manuscripts intended for royalty.
For those fluent in languages like medieval Latin or Old French, abbreviations present an additional hurdle—medieval texts are rife with them. In fact, up to 35-40% of words in 14th-century Latin manuscripts are abbreviated, while Old French shows a lower abbreviation rate of 7 to 12%. In specialized texts, such as medical treatises, half of the letters may be omitted entirely, existing only through inference.
Given these complications, one might wonder why popular AI language models like GPT, Mistral, or Gemini were not utilized. The key issue is that these models function by predicting "tokens," or language-building units, relying on regular patterns that simply do not exist in irregular historical texts. Without stable spelling conventions in medieval French, the potential for generating fictitious words dramatically increases.
Consequently, the development team opted for a shape-based approach to character recognition, focusing on the graphical characteristics of each sign rather than attempting to derive meaning. This method treats every character independently, including accents, which are recognized as separate entities within the transcription process.
This strategy led to the inception of the CATMuS project (Consistent Approaches to Transcribing Manuscripts), spearheaded by Ariane Pinche of CNRS in the early 2020s. The objective was to establish a reliable learning corpus prior to the training of algorithms.
In the course of their work, the team manually transcribed 200,000 lines from 300 manuscripts across 11 languages, representing a range of time from the 9th to the 16th centuries. The chosen documents aimed for maximum diversity rather than adhering to individual researchers' preferences.
Medieval monastic scriptoria were essential hubs for manuscript copying, preserving and disseminating knowledge throughout the ages. These centers of intellect played a critical role in safeguarding classical works while also creating original literature.
The cornerstone principle of the CATMuS project is to leave the text unaltered. Abbreviations were not expanded, spelling was maintained in its original form (as no standardized spelling was available), and even scribal errors, like misplaced letters, were preserved. Nonetheless, some elements, such as abbreviations and strikethroughs, were treated with consistency during transcription.
Using established open-source transcription tools like Kraken and eScriptorium—none of which rely on conventional language models—helped sidestep the notorious problem of model “hallucinations.” While this approach introduced new types of errors, such as misinterpreting "ri" for "n," Clérice expressed his preference for these mistakes over potentially fabricated terms appearing within the text.
The CoMMA initiative applied this innovative model to a vast assortment of previously digitized manuscripts from platforms like the French National Library's Gallica initiative, the online manuscript repository ARCA from CNRS, e-codices in Switzerland, and libraries including Oxford's Bodleian Library and the Bavarian State Library in Munich.
The CoMMA platform presents the transcriptions exactly as generated, without any human oversight or corrections. After evaluating three consecutive lines from 670 manuscripts, the team discovered an average error rate of 9.7%. Each manuscript's metadata also reveals the percentage of lines accurately recognized, with some documents achieving accuracy rates exceeding 80%, while others dip lower. A discernible trend indicates that more cursive handwriting—especially in later manuscripts—leads to poorer recognition results, primarily due to a lack of similar examples in the training dataset.
In total, this project has amassed over three billion words, predominantly in Latin and Old French. The volume of Old French texts available has increased forty-fold, offering unprecedented opportunities for research in areas like historical linguistics, philology, and textual history—fields that previously struggled due to the scarcity of resources.


