The archive · AI & Models · Technical decision · 2026
General Instinct bets frontier models must fit edge devices — 245GB squeezed onto 8GB
Two-person YC P26 startup compresses datacenter-scale frontier models into offline runtimes for robots and drones, open-sourcing its compression engine
General Instinct
What the business is
General Instinct's Instinct Edge takes a frontier model, a target edge device, and a latency budget, and returns an offline runtime tuned to that hardware — no cloud call at runtime.
How it started
Bill Jiao and Guanming Wang, a two-person team in San Francisco, founded General Instinct in 2026 after years of robotics and ML research — Jiao previously worked on Siemens' multimodal foundation model, Wang was a Google DeepMind reinforcement-learning researcher. Their repeated experience: the best models never fit the hardware actually available on robots, drones, and factory devices.
What happened
Instinct Edge distills and quantizes large models to roughly three effective bits, keeping always-active parts such as routers, normalization layers, and vision pathways at higher precision while compressing routed experts aggressively, then uses on-policy distillation to recover lost reasoning. Published results include compressing Qwen3.5-122B-A10B from a roughly 245GB BF16 model to a 48GiB GGUF that runs in 7.6–8GB peak VRAM, and a multimodal classifier on Jetson Orin NX with a 111ms cold start delivering every decision inside a 150ms budget. The compression engine is open source as InstinctRazor.
How it ended up
As of the YesPress profile, General Instinct is still early: two people, YC Spring 2026 backing, and no disclosed customers or revenue. Its Launch HN drew sharp questions from the quantization community about benchmarks against Unsloth and other 3-bit methods, which the founders answered by pointing to reproducible open pipelines.
Background
General Instinct was founded in 2026 by Bill Jiao and Guanming Wang, two robotics and ML researchers who kept facing the same wall: the best AI models are designed for datacenter assumptions, while robots, drones, and edge devices have the opposite constraints. Their bet is that physical AI needs frontier-grade models running on the device itself, fully offline.
The company's Instinct Edge product turns that bet into a pipeline: give it a model, a target chip, and a latency budget, and it returns an offline runtime. Underneath, it quantizes models to roughly three effective bits while protecting always-active components, then uses on-policy distillation to recover capability. The headline result is a 245GB Qwen3.5-122B MoE model compressed to 48GB and runnable in about 8GB of VRAM, with every classifier decision inside a 150ms budget on a Jetson Orin NX.
The two-person startup also open-sourced its compression engine, InstinctRazor, as a deliberate trust and distribution strategy — skeptics can reproduce the claims themselves. General Instinct is betting that as physical AI grows, the integration work of making frontier models fit small chips becomes a market of its own, and that selling that layer beats building another robot.
What has to be true
- Both founders hit the deployment wall personally in robotics, so the company is built for a problem they actually had rather than an adjacent one that demos better
- On-policy distillation plus selective precision quantizes harder than standard 3-bit methods while recovering reasoning, addressing the usual small-and-dim tradeoff
- Open-sourcing InstinctRazor gives skeptical engineers a reproducible proof and seeds adoption among exactly the customers the paid pipeline targets
- Latency, connectivity, and privacy push inference to the device, so a deployment layer for physical AI has a growing addressable market
What can be applied
A wall the founders hit repeatedly is a reliable compass: sell the missing layer that lets any robot run a frontier brain — and open-source the core to earn trust in a hype-heavy market.
Aftermath
As of July 30, 2026, General Instinct remains a two-person, YC-backed San Francisco company with no disclosed customers or revenue. Its public artifacts are the open-source InstinctRazor toolkit, a technical blog on sub-4-bit MoE compression, and the Instinct Edge product targeting Jetson, mobile NPUs, ARM CPUs, Apple Neural Engine, and Snapdragon. Next steps are converting published results into paying physical-AI deployments while defending against tooling like TensorRT and llama.cpp and aggressive community quantization.
Sources
- Launch HN: General Instinct (YC P26) – Frontier models on edge devices
- General Instinct: Frontier AI Models on Any Edge Device
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card