The archive · AI & Models · Product decision · 2026
sllm bets devs split one $14k/mo GPU node: Show HN 188 pts, 104 comments
sllm sells flat-fee slots on shared dedicated GPU nodes for big open models — from $5/month, OpenAI-compatible, no charge until a cohort fills.
sllm
What the business is
sllm pools developers into 'cohorts' that share a dedicated GPU node running open models. Running DeepSeek V3 (685B) needs eight H100 GPUs — about $14,000/month — so the company sells reserved slots from about $5/month for smaller models, spins a node up only when a cohort fills, never charges before then, logs no traffic, and exposes an OpenAI-compatible API through vLLM so users only swap the base URL.
How it started
The founder (HN handle jrandolf) launched sllm on 2026-04-04 with the arithmetic that most developers only need 15–25 tokens per second, while frontier open models like DeepSeek V3 (685B parameters) require eight H100 GPUs at roughly $14,000 per month — capacity a solo developer can never justify. His answer was a timeshare: let a group of developers reserve slots on one dedicated node and each pay a flat monthly fee.
What happened
The Show HN drew 188 points and 104 comments. Commenters pushed on the two weak spots: pricing transparency (the Join button routed to Stripe without showing a price, which the founder admitted was an error and said he would fix) and fairness under shared load (he promised rate limiting and queuing, saying vLLM batches requests continuously with an average time to first token under two seconds). Skeptics compared sllm to vast.ai, TensorDock and OpenRouter; supporters read it as a Kickstarter-style group coupon. Complete AI Training covered the model the next day.
No ending yet — it is still running.
Background
sllm is a startup that lets developers share a dedicated GPU node through a 'cohort'. Running DeepSeek V3 (685B parameters) requires eight H100 GPUs costing roughly $14,000 per month, while most developers only need 15–25 tokens per second, so sllm sells reserved slots from about $5/month for smaller models and only spins up hardware once a cohort fills. No one is charged until then, traffic is not logged, and the service runs vLLM behind an OpenAI-compatible API so existing code needs only a new base URL.
The founding bet is that a flat monthly fee beats metered cloud APIs for developers who want big open models without enterprise budgets. Continuous batching lets one node serve many users simultaneously, and fixed-price 'unlimited' access is the pitch, with rate limiting and queuing promised to keep any single user from hogging the box. The no-charge-until-filled reservation turns each cohort into a group commitment rather than a pay-per-token sale.
The Show HN on 2026-04-04 drew 188 points and 104 comments. The thread's skeptics compared sllm to vast.ai, TensorDock and OpenRouter, questioned whether 'unlimited' survives heavy users, and flagged that the Join button led to Stripe without showing a price — an error the founder admitted and said he would fix. Complete AI Training covered the model the next day, and supporters framed it as a Kickstarter-style group coupon for frontier inference.
What has to be true
- Frontier open models need datacenter hardware most developers can never rent alone, so pooling is the only path from $14k/month to pocket money.
- An OpenAI-compatible vLLM endpoint means adopting sllm is a base-URL change, removing the switching cost that kills new inference providers.
- No-logging privacy gives a flat-fee sharing service a story metered API tiers cannot match for proprietary code and data.
- Not charging until a cohort fills removes the risk of paying for a node that never assembles — but it makes fill velocity the whole company.
What can be applied
Sell the idle capacity, not the API: a $14k/month node becomes $5/month slots, but the bet hangs on cohort fill velocity and fair sharing once strangers share one box.
Aftermath
As of 2026-09-06 the public record shows sllm still on its reservation model: the April 2026 Show HN and the Complete AI Training feature are the milestones, and no later source reviewed for this card reports a filled cohort, a pricing change or funding. The open questions the launch itself exposed — cohort fill velocity, rate-limit fairness under real load and whether fixed-fee 'unlimited' access survives heavy users — remain the difference between a niche timeshare and a business.
Sources
- Show HN: sllm – Split a GPU node with other developers, unlimited tokens
- sllm lets developers split GPU node costs through a cohort sharing model, cutting DeepSeek V3 access from $14,000 to $5 a month
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card