Hot French startup ZML releases free product to speed inference across lots of AI chips
French AI startup ZML has released ZML/LLMD, an inference-server product that lets a range of open-source large language models run across chips from multiple vendors — including Nvidia, AMD, Google TPU, Apple Metal and Intel Arc. The launch matters because inference — the processing of prompts, as opposed to model training — has become a central battleground amid rising fears over AI costs, and ZML's cross-chip approach aims to break vendor lock-in and let companies mix cheaper or more energy-efficient hardware.
Backed by Turing Award winner Yann LeCun and led by founder Steeve Morin (a former Zenly engineering VP), the Paris-based, 20-person firm has raised $20 million from investors including 20VC, Kima Ventures and Kindred Capital. Unlike ZML's earlier open-source framework, LLMD is proprietary but launches free so the company can study usage before deciding on pricing. ZML enters a crowded "inference gold rush" against rivals such as Baseten, Inferact (from vLLM's creators) and RadixArk (behind SGLang), though Morin says his ambitions stretch further, including co-designing silicon and supporting emerging European chipmakers.
- ZML launches free software running open-source LLMs across many chip brands.
- Aims to cut costs and break Nvidia-style vendor lock-in.
- Paris firm has $20m funding and LeCun's backing.
AI Art Business Celebrity Companies Culture Entertainment Software Technology