MiniMax H3 is a video model for creators that turns text, images, video, and audio references into short clips with synchronized sound.
A first-frame generator animates a picture. H3 can also read a collection of reference media and use it to direct a shot. A character photo can define who appears, a video can supply movement, and an audio clip can guide the voice or sound. The prompt explains which parts of those references belong in the result. This makes the model relevant to creators working on recurring characters, product demonstrations, and scenes whose action must follow an existing performance.
The original announcement is dated 31 July 2026. MiniMax published its open-weight announcement on 3 August. These are separate milestones: the hosted product arrived before the public model weights. MiniMax points creators to Hailuo's H3 tool for a browser-based experience.
The official repository separates the system into prompt interpretation, base generation, and 2K regeneration. The distinction matters when choosing between running a model yourself and buying a complete hosted render. Downloading the base checkpoints does not reproduce every component of the hosted system.
The model download on Hugging Face includes separate FL2VA and Ref2VA folders. The repository recommends SGLang, vLLM, diffusers, and ComfyUI workflows. Its SGLang example serves a checkpoint across four GPUs, so local deployment is a developer route requiring suitable hardware rather than a lightweight desktop installation.
H3 Max is fal's post-trained variant of MiniMax H3. fal describes changes aimed at prompt adherence, aesthetics, and throughput. It offers separate text-to-video, image-to-video, and reference-to-video endpoints. The phrase H3 Max Ref refers to that reference workflow, rather than a separate MiniMax product to install.
In fal's reference API, the prompt addresses assets as Image 1, Video 1, or Audio 1. For example, two character pictures can define the subjects while the prompt asks them to walk together through a garden. The API accepts up to 12 reference files in total. Video and audio references are 2-15 seconds per clip, with a combined duration limit of 15 seconds for each modality. Optional first and last frames provide additional control over the beginning and ending.
Use Hailuo for the browser workflow, the MiniMax platform for the official hosted API, or the released base checkpoints for a self-hosted pipeline. fal's Max endpoints are another hosted route with their own input limits and billing. The reference endpoint lists output charges and a separate reference-input token charge, so the length of the finished video alone does not determine its total cost. Check the selected provider's current quote before rendering, especially when uploading video references.