Feedback
AI Ad Video Example
Loading...
minimax h3 video model
Produce 2K clips with stereo audio using the minimax h3 video model — one engine that handles text, images, footage, and sound for up to 15 seconds.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Sora2 AI
Advanced AI Video Generator for High-Quality Videos
What makes the minimax h3 video model a powerful multimodal tool
Built by MiniMax and offered from day one on fal.ai, the minimax h3 video model is an open-weight omni-modal engine that reads text, pictures, footage, and sound in one pass. It outputs 2K clips with native two-channel audio lasting up to 15 seconds, supports targeted scene edits, sharp text and UI rendering, and accepts as many as 12 mixed reference inputs per run.
- A single context for all media typesFeed the minimax h3 video model up to nine pictures, three footage clips, and three audio files together; it merges characters, motion, framing, and sound into a consistent scene.
- Sound that comes with the pictureEach render from the minimax h3 video model includes original music, spoken lines, sound effects, and room tone locked to the cut, plus the ability to transfer or clone voices from supplied audio.
- Surgical edits without ruining the frameSwap objects, change on-screen text, replace voices, or flip daylight to night; the minimax h3 video model updates just the selected area and leaves the surrounding footage untouched.
Getting started with the minimax h3 video model
Follow these three steps to request 2K footage with synced sound from the minimax h3 video model API.
Core capabilities of the minimax h3 video model
From three API endpoints and one shared multimodal context to stereo output, targeted edits, legible typography, and usage-based billing, the minimax h3 video model covers every stage of 2K video creation on fal.ai.
Three ways to start a generation
The minimax h3 video model exposes text-to-video, image-to-video with first/last-frame control, and reference-to-video endpoints, so any creative workflow gets a matching entry point.
Twelve references per request
Bring nine images, three video clips, and three audio tracks into one job; the minimax h3 video model learns characters, style, motion, camera work, composition, and editing rhythm from them.
Clean text and real UI generation
Produce readable captions, end cards, logos, and working interfaces — landing pages, game menus, HUDs, and kinetic type — using the minimax h3 video model.
Long prompts for full-scene control
Put an entire shot list into one request: the minimax h3 video model accepts prompts up to 7,000 characters, giving you total command over every scene.
High-resolution output at 24fps
Generate up to 15 seconds of 2K video with a 1440px short edge, 24 frames per second, six aspect ratios, and an adaptive option through the minimax h3 video model.
Pay only for what you render
The minimax h3 video model uses serverless, usage-based pricing with no subscriptions or minimum commitments, and you retain commercial rights to everything you create.
Common questions about the minimax h3 video model
Quick answers to the most asked questions about running the minimax h3 video model through fal.ai.
What exactly is the minimax h3 video model?
It is an open-weight, general-purpose omni-modal model from MiniMax, available on fal.ai from launch day. A single model ingests text, pictures, footage, and audio together and produces 2K video with native stereo sound for up to 15 seconds.
Which endpoints does the minimax h3 video model expose?
You get three routes: text-to-video, image-to-video with optional first/last-frame control, and reference-to-video that locks subjects, styles, motion, camera moves, and voices to your supplied materials.
What resolutions and lengths are available?
The minimax h3 video model renders 2K footage (1440px on the short side) at 24fps. Duration can be set between 5 and 15 seconds, and supported ratios are 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive.
Does the output include sound?
Yes. Every video created with the minimax h3 video model comes with stereo audio: original music, speech, foley, and ambience synced to the picture, and you can transfer or clone voices from reference audio.
How many images, clips, and audio files can I use at once?
Up to twelve in total — nine reference images, three reference clips of 2–15 seconds, and three audio tracks of 2–15 seconds. Audio must be paired with at least one image or video when using the minimax h3 video model.
Are generated videos ok for commercial use?
Yes. Content created through the fal.ai API with the minimax h3 video model can be used commercially according to fal.ai's terms of service.
Ready to produce video with the minimax h3 video model?
Put the entire minimax h3 video model pipeline to work today — 2K clips, stereo sound, multimodal inputs, precise edits, and flexible API billing on fal.ai.
