Definition◆Article
MiniMax-H3 is a general-purpose, omni-modal generative system that accepts text, images, audio and video inputs and can generate up to 15-second video clips with audio included.
MiniMax recently released H3, an omni-modal model that takes text, images, audio, or video as input and generates video clips with audio. ✦ AI generated
article author · Simon Willison's Weblog · 2026-08-04 · original ↗
MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio included.
Read full article ↗excerpt · fair-use quotation
- ·General-purpose, omni-modal generative system
- ·Accepts text, images, audio, and video inputs
- ·Generates up to 15-second video clips with audio
- ·Supports four distinct input modalities
- ·Unified system for media understanding and creation
Around this claim
This moment responds to
provides context → A Python package has been created to port MiniMax-H3 to MLX for running on Apple Silicon.article author · Simon Willison's Weblogprovides context → Running the model on an M5 Max MacBook Pro downloaded approximately 115 GB of model files and took just under 45 minutes for video generation.article author · Simon Willison's Weblogprovides context → The video output quality is impressive, but without prompt guidance the audio produces nonsensical speech-like sounds.article author · Simon Willison's Weblogprovides context → MiniMax provides a prompting guide with detailed information on how to achieve proper audio generation with the model.article author · Simon Willison's Weblog