Skip to main content
MiniMax H3 video model generating 2K clips, priced at $1.95 for 15 seconds, shown on a screen.

Editorial illustration for MiniMax H3 Video Model Generates 2K Clips, Priced at USD 1.95 for 15 Seconds

MiniMax H3 Generates 2K Video Clips for $1.95

4 min read

MiniMax put a price tag on its newest video model this week: $1.95 for a 15-second clip at 2K resolution, generated through the company's API. The model, called MiniMax H3, went live on July 31, 2026, both as an endpoint under the model ID MiniMax-H3 and inside the consumer-facing Hailuo AI app. Outputs run 4 to 15 seconds, integer durations only, with native stereo audio baked in rather than added afterward.

What separates H3 from the usual video-generation stack is how MiniMax built it. Most competitors ship separate expert models for text-to-video, image-to-video, first-and-last-frame work, subject reference, motion reference, and editing. MiniMax folded all of that into a single pretraining setup, treating text, images, video, and audio as one shared context that the model reads at once. Reference and editing instructions get expressed in plain language instead of routed to a specialized tool.

MiniMax is pitching this at advertising teams, e-commerce sellers, game studios, and film previsualization shops, the kind of users who need product videos, ad variants, or character-consistent cinematics without stitching together five different systems. The API itself runs on an asynchronous three-step flow: create a task, poll for status, download the result.

MiniMax releases MiniMax H3, a general-purpose multimodal generation model. MiniMax H3 is not a text-to-video model with add-ons. MiniMax describes it as a general-purpose multimodal generation model that reads text, images, video, and audio as one unified context and returns video with native stereo sound.

Why this matters

For teams building on video generation, H3's pitch is consolidation: one context window instead of five separate pipelines for text-to-video, image-to-video, and the rest. That's a real workflow change if it holds up, since stitching together subject reference and motion refer tools has been the actual bottleneck for a lot of shops, not model quality. But we'd hold off on the pricing story.

A $1.95-per-15-second figure sourced to third-party trackers, with MiniMax's own pay-as-you-go page still showing only Hailuo 2.3 tiers, is not a confirmed rate card. It's someone's estimate that got repeated. Builders costing out production pipelines should treat $0.13 per second as a placeholder until MiniMax publishes its own numbers, because a rate that moves 20-30% changes the math on whether this replaces existing tools or just supplements them.

Native stereo audio and 2K output are worth testing regardless. Just verify pricing directly with MiniMax before you plan a budget around a number that came from launch coverage rather than the source. Watch for the official pricing page update.

Common Questions Answered

What is the pricing model for MiniMax H3 video generation?

MiniMax H3 is priced at $1.95 for a 15-second clip generated at 2K resolution through the company's API. The model supports output durations ranging from 4 to 15 seconds in integer increments only, making it accessible for various video generation needs.

How does MiniMax H3 differ from traditional text-to-video models?

MiniMax H3 is built as a general-purpose multimodal generation model rather than a simple text-to-video model with add-ons. It reads text, images, video, and audio as one unified context and returns video with native stereo sound, eliminating the need for separate pipelines.

When did MiniMax H3 become available and where can it be accessed?

MiniMax H3 went live on July 31, 2026, and is available both as an API endpoint under the model ID MiniMax-H3 and within the consumer-facing Hailuo AI app. This dual availability allows both developers and general users to access the video generation capabilities.

What workflow advantage does MiniMax H3's unified context provide for development teams?

H3 consolidates multiple separate pipelines into one unified context window, eliminating the need for separate text-to-video, image-to-video, and other specialized tools. This represents a significant workflow change that addresses the actual bottleneck many teams face when stitching together subject reference and motion reference tools.

What audio capabilities does MiniMax H3 include in its video outputs?

MiniMax H3 generates videos with native stereo audio baked directly into the output rather than being added afterward. This integrated approach ensures higher quality audio-video synchronization compared to traditional post-processing methods.

LIVE13:29MiniMax H3 Video Model Generates 2K Clips, Priced at USD 1.95 for 15 Seconds