MiniMax's New H3 Model Generates Video With Real Stereo Sound, and Weights Are Coming
Fifteen-second clips at 2K resolution with dialogue and sound effects generated in the same pass as the picture, not dubbed on after.
MiniMax's new video model generates its own sound instead of dubbing it on afterward.
Shanghai-based MiniMax released H3 on July 31, 2026, a video-generation model that produces up to fifteen-second clips at 2K resolution with native stereo sound, dialogue, effects, and ambience generated in the same pass as the picture rather than added afterward.
What makes the audio different
Most video models generate silent clips that get sound added in a separate pass. H3 generates dialogue, sound effects, and environmental audio together with the video, timed to on-screen action from the start. The model also supports motion transfer between clips and what MiniMax calls omni-reference, multiple reference inputs in a single prompt.
Open weights, but not yet
MiniMax says it plans to release H3's weights within days of the launch. As of this writing that hasn't happened, so treat the open-weights claim as a near-term promise, not something available to download today. MiniMax claims 2K generation at under a third of the cost of mainstream rival products.
The competitive picture
H3 lands in an already crowded Chinese video-model market, positioned against ByteDance's Seedance 2.0 and Kuaishou's Kling 3.0. Open weights, if they ship as promised, would be the clearest differentiator from either competitor, both of which are closed.
Why a build studio cares
A genuinely open-weight video model with usable audio changes the calculus for any client project that needs generated video without a per-call API bill. Worth a second look once the weights actually land.
Next step: read the Reuters report on the launch. If generative video is part of a client pipeline you're evaluating, write to us at hello@gattyworks.com.