Kog is going deeper to squeeze more inference out of GPUs
French startup Kog is developing software intended to make existing data-centre GPUs run AI inference much faster, arguing that better low-level optimisation can reduce a major cost and speed bottleneck without requiring specialised chips. Its early demonstration attracted significant business interest, particularly from companies using AI for professional software engineering and prompt-based app or game creation, where faster results could improve productivity and revenue.
Kog says its technology achieved 3,000 tokens per second for a single request using its open-source, two-billion-parameter Laneformer model on AMD MI300X and Nvidia H200 GPUs, but it must prove the approach works at scale with much larger language models to support its claim of up to 30 times faster inference. Founder Gaël Delalleau believes modern GPUs’ growing memory bandwidth is underused, and says his background in physics and cybersecurity informs Kog’s focus on understanding and optimising GPU behaviour at a very low level.
- Kog aims to accelerate AI inference on existing data-centre GPUs.
- Its early results use a small model, not large commercial LLMs.
- Faster inference could lower costs and improve AI workflow responsiveness.
Americas Art Business Celebrity Companies Culture Entertainment World