reference to video
16 reference to video models and apps on inference shell.
run via API, SDK, or belt CLI.
video
16
grok-imagine-video-1-5
xai/grok-imagine-video-1-5
generate videos with xai's grok imagine video 1.5. text-to-video, image-to-video and reference-to-video at 480p, 720p or 1080p, 1-15 seconds, with generated audio.

MiniMax H3 Max
falai/minimax-h3-max
minimax h3 max — fal's post-trained h3 for stronger prompt adherence and aesthetics. text-to-video, image-to-video with end frame, and reference-based generation. 480p/768p, 5-15s, ~3s per 5s clip.

Seedance 2.5
bytedance/seedance-2-5
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.5 model via byteplus ark api. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

seedance-2-5-studio
bytedance/seedance-2-5-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

PixVerse Fusion
pixverse/fusion
pixverse fusion — generate video from reference images. compose 1-3 subject/background images into a video scene with text-guided motion.

seedance-2-0-studio-mini
bytedance/seedance-2-0-studio-mini
cost-effective multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Mini
bytedance/seedance-2-0-mini
cost-effective multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 mini model via byteplus ark api. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Kling Video O1
klingai/video-o1
kling video o1 (omni) - unified video generation with text, image references, start/end frames, element references, and video references for editing and style transfer. the most capable kling model.

seedance-2-0-studio
bytedance/seedance-2-0-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

seedance-2-0-studio-fast
bytedance/seedance-2-0-studio-fast
fast multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0
bytedance/seedance-2-0
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 model via byteplus ark api. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Fast
bytedance/seedance-2-0-fast
fast multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 fast model via byteplus ark api. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

HappyHorse 1.0 Reference-to-Video
alibaba/happyhorse-1-0-r2v
happyhorse 1.0 reference-to-video generates videos preserving subject characters from up to 9 reference images, with enhanced stability in subject and scene referencing via dashscope api

Wan 2.7 Reference-to-Video
alibaba/wan-2-7-r2v
wan 2.7 reference-to-video generates videos featuring characters from reference images and videos, supporting multi-character interaction, voice timbre cloning, and first-frame control

Grok Reference Video
xai/grok-reference-video
generate videos using reference images for style and content guidance with xai's grok imagine video model. provide reference images to influence the visual style of generated videos.

Veo 3.1
google/veo-3-1
veo 3.1 via vertex ai - advanced video generation with frame interpolation, reference images, and audio generation
explore more on inference shell
reference to video is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.