Bullish

MiniMax Launches H3 Omni-Modal Model with 15s Video and Audio Generation

2026-07-31 10:27:40

MiniMax unveils H3, an omni-modal model generating 15s 2K video with native stereo audio. It consolidates motion, person, and sound reference into one task, priced at $0.13/sec for 2K.

Woofun AI reports that MiniMax has released the H3 omni-modal generation model, which processes text, images, videos, and audio simultaneously to generate or edit video via natural language commands. The model consolidates previously separate tasks, such as motion transfer, person reference, and sound reference, into a single operation, allowing users to combine elements from multiple inputs.

H3 supports video generation up to 15 seconds in 2K resolution with native stereo sound. API pricing is set at $0.13 per second for 2K output, totaling approximately $1.95 for a 15-second clip, while 768P resolution costs $0.09 per second. Each task accepts a maximum of 9 images, 3 videos, and 3 audio clips, capped at 12 files total. MiniMax intends to release model weights within days, though official evaluations against competitors like Seedance and Veo remain pending.

WOOFUN AI

Impact Assessment · Quick Read

The launch of H3 signifies a shift toward unified omni-modal workflows, reducing the complexity and cost of multi-step video production. By integrating motion, visual, and audio references into a single inference step, MiniMax may lower barriers for high-quality video generation. The competitive pricing structure could pressure rivals to adjust their API rates, while the upcoming weight release may spur community-driven optimization and benchmarking.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions