Nvidia's substantial hold on the market is still intact, yet emerging competitors and options are appearing from various quarters.
ZML, an innovative French AI startup backed by Turing Award laureate Yann LeCun, has introduced software designed for inference performance that enables a range of open-source large language models to operate on multiple chipsets, including those from Nvidia, AMD, Google’s TPU, Apple Metal, and Intel Arc.
With the launch of ZML/LLMD, an inference server for large language models, the company aims to dismantle existing barriers and facilitate the use of diverse chips for AI purposes at optimal speeds, or even exceeding their current capabilities, according to ZML founder Steeve Morin's comments to TechCrunch.
As AI technology becomes increasingly embedded in professional and personal spheres, the optimization of inference—the processing of requests—has become more significant than model training, often encountering software and architectural limitations that create vendor lock-in, Morin explained.
Achieving peak performance across different chip technologies represents not only a technological breakthrough but also a potential disruption in the market, especially as concerns about the costs related to AI grow.
ZML aims to provide companies and cloud platforms with the flexibility to employ a combination of chips, some of which may be more budget-friendly or energy-efficient. "The goal is to empower users to construct their own systems and achieve genuine efficiencies that promote the widespread adoption of AI," Morin stated.
This software may also bolster new AI chip innovators, particularly many based in Europe, which Morin noted include firms like Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA. What resonates more with him than their geographical origins is the potential for ZML to collaborate with them on pioneering projects.
Morin, however, does not underestimate Nvidia's market position. He emphasized that ZML maintains a positive relationship with the leading chip manufacturer, which has been preparing for the increased focus on inference.
The surge of investment in inference has been so significant that it’s been dubbed an "inference gold rush." ZML faces competition from notable players such as Baseten, recently valued at $13 billion, Inferact—connected to the creators of the vLLM open-source project—and RadixArk, which is associated with SGLang.
While both vLLM and SGLang are partial competitors to LLMD, Morin's vision for ZML encompasses a wider scope. He noted, "We have reached a stage where we are co-designing silicon." He credited the rapid progress of the Paris startup to its agile team of just 20 members, with numerous future releases on the horizon.
This small team also benefits from substantial financial backing. Morin, who previously served as VP of engineering for Zenly—a company acquired by Snapchat for a hefty sum in 2017—successfully raised $20 million from various venture firms, including notable investors like Harry Stebbings’ 20VC and Xavier Niel’s Kima Ventures.
In contrast to ZML’s earlier public offering, which focused on inference and was launched in 2024 with updates in March, ZML/LLMD is not open source. However, it is being released as a complimentary product to gauge user engagement. “I prefer to assess usage and then monetize in the most impactful ways, avoiding the pitfalls of being overly greedy too early,” Morin remarked.
It's too soon to predict when ZML/LLMD might transition to a paid model, or what user adoption will resemble. Nonetheless, the startup’s capital structure indicates that other founders, such as Solomon Hykes of Dagger and Docker, and Clément Delangue and Julien Chaumond from Hugging Face, are taking notice, as well as LeCun, who is now affiliated with AMI Labs. This trend highlights a growing capability for Europe’s AI startups to thrive locally. “I couldn’t have established ZML anywhere but in Paris,” Morin concluded.



