The archive · Developer & Business Tools · Technical decision · 2025–2026
Kog bets GPU limits are software, not silicon: 30x faster LLM inference on GPUs you own
French solo founder reverse-engineers NVIDIA H200 / AMD MI300X down to assembly; 3,000 TPS demo hits HN front page, 200+ leads
Kog
What the business is
Sells the Kog Inference Engine (KIE), software that speeds up LLM inference on standard datacenter GPUs like the Nvidia H200 and AMD MI300X
Starting capital:Seed round co-led by Varsity VC; backed by France's Bpifrance, the French Tech 2030 program and cloud provider Scaleway (amounts undisclosed)
How it started
Solo founder Gaël Delalleau combines two backgrounds: solid-state physics from École Polytechnique, and four DEFCON CTF finalist appearances as a white-hat hacker. His earlier startup, TC50 2009 alum Stribe, was co-founded with Kamel Zeroual — the Varsity VC founder who later co-led Kog's seed round.
What happened
Kog emerged from stealth in May 2026 with a Hacker News tech preview claiming 'extremely fast single-request decoding' on the AMD MI300X and Nvidia H200 — 3,000 per-request tokens per second with Laneformer 2B, since open-sourced. The demo generated 200+ business leads. Listening to prospects, Kog learned enterprise customers wouldn't fine-tune small models, so it refocused from small-model demos to accelerating the large models enterprises actually deploy.
How it ended up
Still pre-product-market fit as of Sep 2026: an 11-person team hand-optimizing chip by chip, working toward a first major model at 10x speed in September 2026 — the milestone that is supposed to unlock a Series A.
Background
Kog is a bet against the direction of the entire AI inference industry. While markets celebrated Cerebras' IPO for purpose-built inference chips, French startup Kog claimed the real headroom sits inside the GPUs enterprises already own. Its May 2026 Hacker News tech preview demonstrated 3,000 tokens per second per request on the AMD MI300X and Nvidia H200 — standard datacenter hardware — using a purpose-built 2-billion-parameter model it then open-sourced as Laneformer 2B.
The method comes from the founder's unusual double training. Gaël Delalleau studied solid-state physics at École Polytechnique and competed four times in DEFCON's Capture The Flag tournament. He describes Kog's approach as understanding 'the laws of physics, and the laws of the GPU' — reverse-engineering hardware down to assembly and binary code and using it for a purpose it wasn't designed for. That means weeks or months of manual research per chip, a constraint that shapes the whole company: with 11 people, Kog can only support a few architectures at a time.
The positioning matters in a crowded field. ZML, also French, bypasses CUDA with a hardware-agnostic software layer; Kog explicitly distinguishes itself, comparing itself to Stanford's Hazy Research with an even deeper focus on GPU acceleration. Its wedge is that decode speed — the sequential part of generation — is bottlenecked by software, not silicon, and that newer GPUs carry memory bandwidth that the vendor stack never unlocks. The demo validated that enough to generate more than 200 tangible business leads.
The market told Kog to change targets. Prospects were unwilling to fine-tune small specialized models, so the startup pivoted its engineering focus from its 2B demo model to accelerating the large models businesses already run. That is a far harder technical target, and Delalleau concedes the company must prove the approach transfers to LLMs before it can raise a Series A. European sovereignty tailwinds — Bpifrance, French Tech 2030, Scaleway — keep the company funded while it tries.
What has to be true
- Delalleau inverted the industry question: instead of 'what chip do we need', he asked 'what are we leaving on the table in chips we already own' — a cheaper bet than fabricating silicon
- The hacker's method is the moat: assembly-and-binary-level reverse-engineering of each GPU takes months, which is hard to copy and hard to automate — and Kog can only support a few chips because of it
- Demos on real enterprise hardware (H200, MI300X) rather than benchmarks on exotic silicon made the claim credible to the exact buyers with 200+ leads
- Pivoting to large models in response to prospect behavior, not founder preference, kept the company aligned with where enterprise money actually sits
What can be applied
When everyone responds to a hardware bottleneck by buying new hardware, the contrarian read is the hardware you own is underprogrammed — but only if you're willing to go deep enough to prove it
Aftermath
As of 2026-09-03 Kog is a pre-revenue, 11-person Paris startup facing a self-imposed deadline: a first major model at 10x speed by September 2026, which CEO Delalleau says will unlock a Series A. Laneformer 2B is open-sourced; design partners are building prompt-to-app products on KIE. The team plans to encode its manual chip-optimization into agentic pipelines. Bpifrance, French Tech 2030 and Scaleway give sovereign-tech tailwinds; Varsity VC is the disclosed lead. The open question: does the 30x speedup generalize from a 2B model to frontier-scale LLMs?
Sources
- Kog is going deeper to squeeze more inference out of GPUs
- French Startup Kog Targets 30x Faster GPU Inference Speeds
- Hacker News: Kog — real-time LLM inference on standard GPUs
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card