MiniMax Launches H3 Omni-Modal Model with 15s Video and Audio Generation
MiniMax unveils H3, an omni-modal model generating 15s 2K video with native stereo audio. It consolidates motion, person, and sound reference into one task, priced at $0.13/sec for 2K.
Woofun AI reports that MiniMax has released the H3 omni-modal generation model, which processes text, images, videos, and audio simultaneously to generate or edit video via natural language commands. The model consolidates previously separate tasks, such as motion transfer, person reference, and sound reference, into a single operation, allowing users to combine elements from multiple inputs.
H3 supports video generation up to 15 seconds in 2K resolution with native stereo sound. API pricing is set at $0.13 per second for 2K output, totaling approximately $1.95 for a 15-second clip, while 768P resolution costs $0.09 per second. Each task accepts a maximum of 9 images, 3 videos, and 3 audio clips, capped at 12 files total. MiniMax intends to release model weights within days, though official evaluations against competitors like Seedance and Veo remain pending.
Comments
No comments yet.