When sales representatives from major data-labeling firms in Silicon Valley attended the International Conference on Machine Learning in Seoul this year, they were geared up to attract significant clients within the AI sector. These companies, collectively valued at tens of billions and generating substantial profits by supplying training data to organizations such as OpenAI and Anthropic, are accustomed to navigating the demands of the often unpredictable AI labs.
However, they discovered another potential market thriving on their doorstep: China’s rapidly advancing AI sector.
Several Chinese firms have laid out their acquisition needs. Tencent, a corporation identified by the U.S. as having ties to the military—an assertion they contest—distributed a comprehensive request for training data that includes areas such as finance and cybersecurity, as well as ambitious objectives like developing AI systems that can enhance their own learning processes.
While the U.S. has long restricted China's access to the advanced chips that drive premier AI models, it has not imposed similar limitations on the data necessary for training these systems. This expertly curated human knowledge, developed through specialized networks and shaped by detailed task specifications, serves as the backbone of the data infrastructure that trains models to handle sophisticated tasks, such as financial analysis and software development. This lucrative industry represents hundreds of millions of dollars each year and is contributing to the relentless pursuit of technological parity between Chinese and American AI models.
Nathan Lambert, previously a research scientist at the Allen Institute for AI, stated, “After the compute needed to train the models, quality data is paramount. As we aim for models to tackle new and more complex areas, securing high-quality data becomes the most impactful factor in making a model effective.”
Forbes has reviewed communications indicating a burgeoning trade worth several hundred million dollars in training datasets between American data firms and China's foremost AI laboratories. The same startups servicing OpenAI and Anthropic are also supporting top Chinese labs, effectively putting Silicon Valley in a position where it is fueling both sides of the evolving AI rivalry.
In the quest to match American AI capabilities, Chinese labs can pursue three main strategies: recruiting top talent from the U.S., distilling insights from outputs generated by systems like ChatGPT or Claude (which technically contradicts their terms of service), and more straightforwardly, acquiring the same training data from the same suppliers as American labs.
Training an AI to master intricate tasks, such as complex accounting or experimental protocols, goes beyond mere data collection. The real advantage lies in the design of the data itself, including which data types to emphasize and the sophisticated criteria crafted by experts to guide the model’s learning process. By purchasing these proprietary datasets, Chinese labs efficiently acquire the extensive expertise that fuels Silicon Valley’s advanced models.
Sean Cai, an AI data consultant and blogger, summed up the situation: “Analyzing the data landscape in China and the overall supply chain reveals that a considerable amount of U.S. data facilitating advancements in the U.S. is also being sold to Chinese labs.”
The shadowy data supply chain in China
This trade network channels directly from American data-labelling companies to leading Chinese technology giants like Tencent. In one correspondence reviewed by Forbes, a Tencent data procurement officer mentioned that they partner with Surge AI, which has provided data to the likes of the U.S. Army and Air Force. Additionally, Mercor—recently awarded a contract with the U.S. federal government—also collaborates with Tencent.
Further correspondence exhibits a similar pattern. Executives from Ant Group, the tech and finance entity behind Alipay, acknowledged a strong partnership with AfterQuery, a U.S. AI training data firm. Conversations with representatives from Alibaba—China’s e-commerce powerhouse—pointed to established contracts with both AfterQuery and Mercor. There’s also evidence that Turing, another data labeling startup from the U.S., has done business with ByteDance, TikTok's parent company.
None of the Chinese companies, including ByteDance, Alibaba, Tencent, or Ant Group, responded to requests for comments. Mercor chose not to provide feedback, but AfterQuery and Surge both maintain they do not disclose customer information. Turing mentioned in a statement: “We are at the forefront of AI development, encompassing both proprietary and open-source models. We anticipate growth in both sectors, with some of today's leading open-source models emerging from non-U.S. labs.”
Collectively, the leading six Chinese AI labs are estimated to spend around $500 million annually on American data labeling services, as reported by two entrepreneurs in the sector citing insights from Tencent and ByteDance executives. Observations from industry insiders indicate that these laboratories see themselves not as direct competitors but as a collective seeking high-quality Western datasets.
American data labeling firms appear willing to sell to this market. Reports suggest that Mercor’s revenue from Chinese labs accounted for about 2% of its total income (with annual revenue surpassing $2 billion by June). At the same time, a significant chunk of AfterQuery's revenue, at least $50 million, arises from Chinese AI companies. Surge AI has been proactive in approaching these clients, with CEO Edwin Chen even visiting China to forge relationships.
Data providers offer bespoke, tailored projects to Chinese AI labs, along with what are termed "off-the-shelf" (OTS) datasets. These pre-packaged training sets are designed for resale to various labs. While this might seem harmless, off-the-shelf datasets are quite powerful; they are developed using robust "knowledge pipelines" created during custom engagements with firms like OpenAI or Anthropic. Consequently, Chinese companies that purchase these datasets gain access to the expertise, rigorous quality checks, and evaluation criteria that have contributed to the training programs of major U.S. AI labs, drastically reducing their developmental timeline.
Cai emphasized the significance of off-the-shelf datasets: “The structure of this data frequently originates from the same vendors who shaped initial scaling strategies for firms like Anthropic,” he noted. “It’s incredibly straightforward for a vendor that has worked with Anthropic to generate similar datasets for a Chinese lab.”
The lucrative nature of off-the-shelf datasets—due to their ability to be resold to multiple clients—generates notable profit margins. Chinese AI labs distinctly express to American data suppliers their desire to procure all data that U.S. labs have already acquired, according to an executive familiar with interactions with these labs.
Max Gazor, founder of Striker VC, which has invested in numerous American AI ventures, acknowledged the ethical complexities surrounding this marketplace. “It's fundamentally an ethical choice regarding whether a company should sell data to Chinese AI outfits,” he remarked.
The regulatory grey area
The U.S. has devoted considerable effort to building barriers around advanced semiconductor technology. However, the vital human expertise required for training AI continues to flow across borders with little oversight.
This situation has bolstered the growth of Silicon Valley’s data companies, many of which are now valued in the billions. Mercor, founded in 2020, is aiming for a valuation near $20 billion, while Surge AI is estimated to exceed $25 billion. Additionally, AfterQuery has seen its revenue soar from $100 million to hundreds of millions in just a few months. Although most of their earnings still come from established U.S. businesses, the appeal of international sales as a lucrative revenue source is undeniable.
There are concerns among some commentators that selling datasets to Chinese firms poses risks to national security. “Some data firms are collaborating with foreign adversaries, and the repercussions are evident,” noted Ali Ansari, CEO of Micro1. “It’s hypocritical to assert American AI leadership while selling millions worth of data to nations we consider adversaries.”
Micro1 asserts that it refrains from selling data internationally, whereas Scale AI had previously explored a contract with ByteDance but withdrew in 2024 due to security apprehensions.
Conversely, proponents of data sales argue that restricting these transactions could stifle international open-source innovation.
“Effectively limiting data exports or AI transfers could damage American open-source efforts,” Cai pointed out. “Such actions would potentially give an unfair advantage to OpenAI and Anthropic, which comes with its own set of challenges.”




