AI Independence Arrives at the Company: Insights from Thomson Reuters’ $40 Million Model

AI Independence Arrives at the Company: Insights from Thomson Reuters’ $40 Million Model
Summary
Thomson Reuters launched its proprietary AI model, Thomson, emphasizing control over proprietary content.
The development of Thomson involved a $40 million investment and collaboration with experts.
Companies can enhance sovereignty by managing their unique data and reducing reliance on external models.

Share

Bookmark

Newsletter

In November, I suggested that artificial intelligence has become as essential as steel and oil, highlighting the necessity for nations to develop what could be termed "sovereign AI." This refers to a country’s ability to create its own intelligence tools rather than depending on external sources. Since then, many business leaders have asked about the implications of AI sovereignty for their own organizations.

A significant response came on August 24 from Thomson Reuters, which unveiled Thomson, its inaugural proprietary large language model. Rather than crafting a model entirely from the ground up, Thomson Reuters utilized a robust open-weight foundation and enriched it with its proprietary content, expert input, and innovative training methodologies. They are also considering options to host the model within a client's controlled cloud environment, allowing direct interaction with the client's intellectual property while safeguarding it from external providers.

The financial aspect of this endeavor merits attention. Thomson Reuters reported an approximate investment of $40 million in developing Thomson, which included the acquisition of the startup Safe Sign Technologies. The final training phase incurred an estimated GPU expense of under $450,000. However, this figure doesn't encapsulate the full financial picture, as it excludes costs related to the acquisition, staffing, infrastructure, data preparation, and ongoing operational expenses.

In terms of performance, early results are noteworthy, although they require nuance. The technical report indicated that Thomson-1.0-Large outperformed both GPT-5.4 and Claude Sonnet 5 in aggregate benchmark scoring but fell short against Claude Opus 4.8 in certain areas, such as coding and abstract reasoning.

Thomson Reuters also conducted a blind comparison study involving 35 attorney-editors and more than 3,000 evaluations. The study generally favored the Thomson system over OpenAI and Anthropic offerings, particularly in the legal domain. However, it should be noted that this comparison assessed entire systems rather than just models, as Thomson had the advantage of accessing Westlaw, Practical Law, and Reuters, whereas the competing systems were limited to web search capabilities.

Despite the qualifications, the findings are significant. Thomson Reuters didn't simply rely on a superior model; rather, it built an effective intelligence generation system tailored to its expertise.

Thomson Reuters' approach involved acquiring Safe Sign, a startup with a talented team drawn from prestigious institutions like Cambridge, DeepMind, and Harvard. Following the acquisition, the company established a Frontier AI Research Lab in partnership with Imperial College London. The full version of Thomson utilized Alibaba's Qwen3.5-397B open weights and was refined through extensive cooperation with Imperial, along with internal sources like Westlaw and Practical Law. This effort was supported by a relatively small team, reported to be around three dozen engineers and scientists, employing up to 368 Nvidia B200 GPUs.

Moreover, the development heavily relied on input from hundreds of subject-matter experts who helped shape the training and evaluation processes. Thomson Reuters curated over 11,000 evaluation items, which involved significant investment of attorney time to ensure answers were not only accurate but also relevant and useful for professional needs.

Two critical aspects deserve every CEO's attention. Firstly, Thomson Reuters regularly updated the core model to align with advancements in the open-weight framework, focusing its investment on specialized improvements rather than replicating earlier training stages. This strategy allows for adaptability, although still involves significant resources for retraining and evaluation. Secondly, the company disclosed that it has utilized less than 10% of its legal content in the Thomson model to date, pointing to considerable untapped potential.

In a structured AI sovereignty framework I proposed in November, I identified five crucial layers: energy, hardware, data, models, and talent. A comprehensive sovereign strategy doesn’t necessitate ownership of all elements but emphasizes understanding where dependencies can exist and where control is vital. Thomson Reuters made conscious decisions in this regard, opting to rent computing resources while maintaining control over the aspects that set it apart, including proprietary data and expert evaluations. The company’s technical report aptly notes that sovereignty exists on a spectrum; Thomson Reuters has reconfigured its dependencies rather than eliminating them.

This perspective aligns with the idea of a “third way” strategy I discussed regarding countries like France and Singapore. An organization doesn't have to match the capital spending of tech giants like OpenAI or Anthropic; it must be strategic about the layers that are becoming commoditized while safeguarding its unique assets. Thomson Reuters is even exploring options for integrating Thomson directly into client environments, providing companies more authority over their proprietary knowledge.

There are three compelling reasons why companies should consider this model, listed in order of urgency. Firstly, proprietary data can be a strategic asset—but only if it is effectively utilized. Organizations in sectors like insurance and healthcare often have vast archives of valuable, proprietary materials that must be correctly organized and linked to experts for effective model training. Secondly, while the economics of AI have evolved, they haven’t vanished; significant capital is still needed to develop a foundation model, though ongoing training of a strong open-weight model can be more economical.

Lastly, dependency should be viewed as a chosen strategy. A company relying solely on external AI solutions accepts the structure, timeline, and terms dictated by those providers. Developing or owning a specialized model can mitigate that risk and significantly reduce costs while keeping sensitive data secure. However, it’s crucial to remember that ownership doesn’t eliminate all dependencies. For instance, Thomson still relies on cloud infrastructure and needs to continually update its systems. In this landscape, having the flexibility to choose how to deploy various models, including Thomson and others, may represent the highest form of sovereignty.

Loading comments...