apps/falai/minimax-h3-max

minimax-h3-max

MiniMax H3 Max — fal's post-trained H3 for stronger prompt adherence and aesthetics. Text-to-video, image-to-video with end frame, and reference-based generation. 480P/768P, 5-15s, ~3s per 5s clip.

run with your agent
# install belt
$curl -fsSL https://cli.inference.sh | sh
# view schema & details
$belt app get falai/minimax-h3-max
# run
$belt app run falai/minimax-h3-max

api reference

about

minimax h3 max — fal's post-trained h3 for stronger prompt adherence and aesthetics. text-to-video, image-to-video with end frame, and reference-based generation. 480p/768p, 5-15s, ~3s per 5s clip.

1. calling the api

install the client

the client provides a convenient way to interact with the api.

bash
1pip install inferencesh

setup your api key

set INFERENCE_API_KEY as an environment variable. get your key from settings → api keys.

bash
1export INFERENCE_API_KEY="inf_your_key"

run and get result

submit a request and wait for the final result. best for batch processing or when you don't need progress updates.

python
1from inferencesh import inference23client = inference()456result = client.run({7        "app": "falai/minimax-h3-max",8        "input": {}9    })1011print(result["output"])

stream live updates

get real-time progress updates as the task runs. ideal for showing progress bars, partial results, or long-running tasks.

python
1from inferencesh import inference23client = inference()456# stream=True yields updates as they arrive7for update in client.run({8        "app": "falai/minimax-h3-max",9        "input": {}10    }, stream=True):11    if update.get("progress"):12        print(f"progress: {update['progress']}%")13    if update.get("output"):14        print(f"output: {update['output']}")

2. authentication

the api uses api keys for authentication. see the authentication docs for detailed setup instructions.

3. files

file inputs are automatically handled by the sdk. you can pass local paths, urls, or base64 data.

automatic upload

the python sdk automatically detects local file paths and uploads them. urls are passed through as-is.

python
1# local file paths are automatically uploaded2result = client.run({3    "app": "falai/minimax-h3-max",4    "input": {5        "image": "/path/to/local/image.png",  # detected & uploaded6        "audio": "https://example.com/audio.mp3",  # url passed through7    }8})

manual upload

you can also upload files manually and use the returned url.

python
1# upload and get a hosted URL2file = client.files.upload("/path/to/file.png")3print(file.uri)  # https://cloud.inference.sh/...

4. webhooks

get notified when a task completes by providing a webhook url. when the task reaches a terminal state (completed, failed, or cancelled), a POST request is sent to your url with the task result.

python
1result = client.run({2    "app": "falai/minimax-h3-max",3    "input": {},4    "webhook": "https://your-server.com/webhook"5}, wait=False)

webhook payload

your endpoint receives a JSON POST with the task result:

json
1{2  "id": "task_abc123",3  "status": 9,4  "output": { ... },5  "error": "",6  "session_id": null,7  "created_at": "2024-01-15T10:30:00Z",8  "updated_at": "2024-01-15T10:30:05Z"9}
idstringtask id
statusnumberterminal status (9=completed, 10=failed, 11=cancelled)
outputobjecttask output (when completed)
errorstringerror message (when failed)
session_idstringsession id (if using sessions)
created_atstringiso timestamp
updated_atstringiso timestamp

5. schema

input

promptstring*

video prompt. max 50000 chars. use timeline markers like [0s-3s] for scene pacing.

example: "A cinematic tracking shot through a misty forest at dawn, golden light filtering through ancient trees."
imagestring(file)

first frame image for image-to-video. formats: jpg, png, webp, gif, avif.

end_imagestring(file)

last frame image. requires image to be set as first frame.

reference_imagesarray

reference images for style/subject guidance (max 9). cannot combine with first/last frame.

reference_videosarray

reference videos for motion/style (max 3). mp4, 2-15s each.

reference_audiosarray

reference audio clips (max 3). requires reference image or video. 2-15s each.

durationinteger

video duration in seconds (5-15).

default: 5min:5max:15
resolutionstring

output resolution. 768p for best quality, 480p for faster/cheaper.

default: "768P"
options:"480P""768P"
aspect_ratiostring

aspect ratio. text-to-video requires explicit value (not adaptive). reference mode defaults to adaptive.

default: "16:9"
options:"adaptive""21:9""16:9""4:3""1:1""3:4""9:16"
prompt_expansion_modestring

balanced (~1s) or quality (~30s extra) for richer prompt rewriting.

default: "balanced"
options:"balanced""quality"
seedinteger

random seed for reproducible generation.

output

expanded_promptstring

the prompt after expansion, as sent to the model.

seedinteger

the seed used for generation.

videostring(file)*

the generated video file.

ready to run minimax-h3-max?

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.