ATRIUMsearch → argument graph
DefinitionArticle

MiniMax-H3 is a general-purpose, omni-modal generative system that accepts text, images, audio and video inputs and can generate up to 15-second video clips with audio included.

MiniMax recently released H3, an omni-modal model that takes text, images, audio, or video as input and generates video clips with audio. ✦ AI generated

article author · Simon Willison's Weblog · 2026-08-04 · original ↗

MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio included.

Read full article ↗excerpt · fair-use quotation

Around this claim