EN
Back to the archive

The archive · AI & Models · Product decision · 2026

MiniMax H3: open 33B video-model weights, sell the Context-IR API

MiniMax open-sourced its 33B H3-Base weights while keeping Context-IR and 2K regeneration hosted; the GitHub repo passed 8.0k stars in five weeks.

MiniMax Group (稀宇科技)

The betThat giving away the 33B H3-Base weights wins developer adoption while Context-IR orchestration and 2K regeneration stay hosted — the only route to official-grade output.Live

What the business is

MiniMax, the Chinese AI company behind Hailuo video generation, sells omni-modal generation through web and desktop apps (hailuoai.video, hub.minimax.io) and a pay-as-you-go API (platform.minimax.io); the H3 release adds open 768p base weights, agent prompt skills and third-party framework integrations, with 2K output and context-rewrite quality gated behind hosted endpoints.

How it started

MiniMax released H3 on 2026-07-30 as a general-purpose omni-modal system that understands text, images, video and audio in one context and generates video with native stereo audio (32 kHz, 4-15 seconds, 24 fps). The open-source release covered H3-Base only: two CFG-distilled checkpoints, FL2VA (text, first/last-frame) and Ref2VA (up to 9 reference images, 3 video clips, 3 audio clips), built around a 33B single-stream Omni-Transformer with roughly 13B AdaLN parameters, plus nine bundled agent skills, one of them (h3-prompt-writing) portable across Claude Code, Cursor, Windsurf, LangChain and OpenAI-based agents. H3-Context-IR, the hosted orchestration system that parses complex multimodal instructions, was explicitly not included in the release, and H3-Regenerate-2K, which regenerates 768p output at 2K in context, was API-only from day one.

What happened

MiniMax positioned the release as an ecosystem move rather than a giveaway: the README ships deployment recipes for SGLang and vLLM, a ModularPipeline integration for Hugging Face diffusers, and official ComfyUI tutorials plus downloadable T2V/I2V/R2V workflow templates. Its recommended 'Full 2K Workflow' combines a locally deployed SGLang H3-Base with the hosted video-generation-v2-h3-context-ir and video-generation-v2-regeneration API endpoints, and ComfyUI's tutorial states that commercial use of locally generated output requires a MiniMax commercial license sold through Comfy, the only official reseller. By the 2026-09-04 crawl the repo had reached 8.0k stars, 545 forks, 64 watchers and 42 commits — about five weeks after creation.

No ending yet — it is still running.

Background

MiniMax H3 is a general-purpose omni-modal system that understands text, images, video and audio in one context and generates video with native stereo audio (32 kHz, 4-15 seconds). On 2026-07-30 MiniMax open-sourced the 768p H3-Base: two CFG-distilled checkpoints (FL2VA for text and first/last-frame modes; Ref2VA for up to 9 images, 3 video clips and 3 audio references), a 33B single-stream Omni-Transformer, and nine agent skills — including h3-prompt-writing, portable across Claude Code, Cursor, Windsurf, LangChain and OpenAI-based agents.

The bet was a split release: give away the model the community can run locally, keep the quality-defining parts hosted. H3-Context-IR, the preprocessing system that refines complex multimodal instructions, was explicitly not in the open-source release; H3-Regenerate-2K, which lifts 768p results to 2K, was API-only. Official SGLang, vLLM, diffusers and ComfyUI recipes lowered switching costs, while the README's 'Full 2K Workflow' pairs local H3-Base with the hosted context-IR and regeneration endpoints — making the API the only route to official-grade output.

Distribution worked fast: the repo was created 2026-07-30 and by the 2026-09-04 crawl showed 8.0k stars, 545 forks, 64 watchers and 42 commits. ComfyUI ships native MiniMax-H3 nodes and notes that commercial use of locally generated output needs a MiniMax commercial license sold through Comfy as the only official reseller. As of 2026-09-05 the release remains live with no revenue or usage figures visible in the material — the monetization leg of the bet is still unproven.

What has to be true

  • The split answered the market's real question: developers could verify H3 on open 768p weights, while hosted Context-IR and 2K regeneration kept differentiated output one API call away.
  • Shipping official SGLang, vLLM, diffusers and ComfyUI recipes made H3 nearly free to adopt, so third-party frameworks did the integration work and turned the repo into the obvious open-video baseline.
  • The bundled h3-prompt-writing skill and publishing guides carried MiniMax's prompting method into Claude Code, Cursor and OpenAI agents, spreading the exact input format its API expects.
  • Keeping Context-IR and Regenerate-2K closed protected the quality gap (faithful complex instructions, 2K regeneration) that weight clones could not reproduce from the base checkpoint alone.
  • Five weeks to 8.0k stars proves the adoption half of the bet, but no funding event, revenue or usage figures appear in the material — monetization remains the unproven half.

What can be applied

Open what the ecosystem can run and keep what decides quality: MiniMax gave away 33B of weights, then monetized the orchestration layer that turns prompts into reliable high-resolution output.

Aftermath

As of 2026-09-05 MiniMax H3 is live: the GitHub repo holds 8.0k stars, 545 forks and 64 watchers; SGLang, vLLM, Hugging Face diffusers and ComfyUI all ship native MiniMax-H3 support; the Hailuo web app exposes the model to consumers; and platform.minimax.io sells pay-as-you-go access to MiniMax-H3 alongside a faster MiniMax-H3-Max variant. The material shows no revenue or usage figures for the release, so the open-weights-to-API conversion story is still being written.

Sources

spotted an error? The archive wants to know.

Your turn

You just read one. Describe what you are building, and see who is betting on the same thing.

Free account · 3 free questions · no card

Related cases