Editorial illustration for China's MiniMax H3 Tops AI Video Ranking With Multimodal Model
MiniMax H3 Tops Video AI Rankings With Multimodal Model
Open-weight video models have spent the past year chasing the closed ones. MiniMax just closed the gap, at least on paper. The Shanghai-based company released H3, a 33-billion-parameter model that handles text, images, video, and audio in one system, and Artificial Analysis now ranks it first in Video Editing among all models tested, open or closed.
That's a first for an open release on this particular leaderboard. H3 also lands second in Text-to-Video and third in Image-to-Video, putting it in range of the proprietary systems that have dominated those categories. The model generates clips between four and 15 seconds with stereo sound, and its model card describes prompts that can combine up to nine reference images, three video clips, and three audio clips at once.
MiniMax didn't open everything, though. Some pieces of the pipeline stay behind closed doors, and the license comes with a revenue cap that limits who can use it commercially. The timing is also notable: ByteDance dropped its own closed model, Seedance 2.5, on the same day, with longer clips and built-in audio baked in.
MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video.
Why this matters
MiniMax just handed developers a 33-billion-parameter model that beats or matches closed rivals on Artificial Analysis's own leaderboard, and it published the weights. That's the part worth sitting with. Sora 2 and Veo 3 have set the bar for video generation over the past year, but both stay locked behind APIs and usage limits.
H3 doesn't. Anyone can download it, fine-tune it, run it on their own hardware, and build products without asking Google or OpenAI for permission. For founders building video tools, that changes the cost calculus overnight.
For researchers, it means an actual object to study rather than a black box behind a rate limiter. We'd still want independent verification beyond one ranking site before crowning anything, and open weights don't guarantee open training data or safety documentation. But a Chinese lab shipping a top-three multimodal video model, in the open, with stereo audio and nine-image prompting, is a concrete signal that the gap between open and closed video generation is narrowing faster than most roadmaps assumed.
Common Questions Answered
What makes MiniMax H3 unique compared to other open-weight video models?
MiniMax H3 is a 33-billion-parameter multimodal model that handles text, images, video, and audio all within one system. It is the first open-weight model to rank first on Artificial Analysis's Video Editing leaderboard, closing the gap that open-weight models have been chasing against closed models for the past year.
How does H3 rank across different video AI tasks on Artificial Analysis?
According to Artificial Analysis, H3 ranks first in Video Editing, second in Text-to-Video generation, and third in Image-to-Video generation. This multi-category performance demonstrates H3's versatility across various video-related AI tasks.
What is the key advantage of MiniMax releasing H3's weights compared to Sora 2 and Veo 3?
Unlike Sora 2 and Veo 3, which remain locked behind APIs and usage limits from OpenAI and Google, MiniMax published H3's weights openly. This allows developers to download the model, fine-tune it, run it on their own hardware, and build products without needing permission or paying for API access.
Why is MiniMax H3's open release significant for the AI development community?
H3's release represents a major milestone where an open-weight model has achieved top performance on a major AI video ranking for the first time. The availability of the model weights enables developers to have full control over deployment, customization, and product development without relying on closed proprietary systems.
Further Reading
- China's MiniMax releases H3 video model - Reuters
- MiniMax challenges ByteDance with low price, open weights for new H3 model - South China Morning Post
- MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities - MiniMax
- MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio - MarkTechPost
- MiniMax H3 - Open-Weights General-Purpose Multimodal Video Model - Fal.ai