In an intriguing twist of irony involving artificial intelligence (AI), major tech companies are facing challenges they had long overlooked. For years, industry leaders contended that if data exists online, it’s fair game for developing AI models. This stance has clashed with content creators striving to protect their intellectual property, often with little success.
Anthropic, OpenAI, and Google now find themselves grappling with the very issues that others have faced as a consequence of putting content on the internet. Specifically, they are dealing with a method known as "distillation," which involves leveraging outputs from one AI model to enhance another. Anthropic has raised alarms that competitors are capitalizing on its research, effectively converting billions invested into AI breakthroughs into shortcuts for other firms. Similar concerns have been echoed by OpenAI and Google.
The challenge is clear: why invest significantly in creating advanced AI models when competitors can replicate much of that knowledge at a fraction of the cost?
What complicates this situation is a striking paradox. From a broader perspective, the process of distillation resembles what AI companies have historically done with internet content: scraping data without consent, transforming it into marketable products, and arguing that it falls under fair use, all while leaving legal nuances for later resolution.
Anthropic maintains that its algorithms are being exploited by rivals, while website administrators have been claiming for years that Anthropic has been draining insights from their content. This reciprocal accusation raises questions about compliance with terms of service from both sides.
Despite branding itself as a leader in ethical AI development, Anthropic stands as one of the more controversial players. Its algorithms aggressively access web content multiple times, yielding minimal return traffic to those sites.
Moreover, while Anthropic, OpenAI, and Google frame the problem as a security breach—citing ongoing bot "attacks" designed to extract intelligence from their models—these very companies have long engaged in similar practices, inundating various websites with bot activities that inflate operational costs. Webmasters find themselves in the uncomfortable position of having their content used without approval, all while they shoulder increased financial burdens.
AI experts distinguish between distillation and web scraping, but there’s still no consensus within the AI sector regarding the ethical boundaries of distillation. Traditionally, distillation may refer to using outputs from a company’s own models to create smaller, more efficient models. In contrast, what Anthropic labels "distillation attacks" refers to situations where competitors use outputs from other companies to refine their AI technologies.
The ambiguity in definitions adds to the confusion; some researchers worry that Anthropic's hardline stance could inadvertently jeopardize all forms of distillation. Nathan Lambert, an authority in open-source AI, describes this escalating anxiety as "distillation panic."
In a nutshell, the major AI firms maintain that utilizing information from the internet without permission is perfectly acceptable, yet when it comes to distillation involving their proprietary content, they assert that such practices cross the line.
This convoluted argument is being undermined by the harsh realities of the digital landscape. Anthropic has limited access to its most advanced models in an attempt to thwart competitors from extracting valuable insights. However, these restrictions have either failed or only spurred the creation of new, clever ways to circumvent them.
Once data is shared online, innovators will inevitably find means to gather, repurpose, and profit from it—applying equally to blogs, photos, software, videos, and the AI outputs that these companies hold dear.
As Zilan Qian from the Oxford China Policy Lab aptly noted, the landscape is akin to a "cat-and-mouse game." As long as AI model outputs are publicly available, individuals will likely devise methods to access them.
It’s possible that distilling another entity’s AI model could even align with fair use principles—a fact that presents a double-edged sword in legal contexts.
Welcome to the current state of the online ecosystem, Anthropic, OpenAI, and Google. It's wise to acclimatize to this evolving environment.


