apps/text to video

text to video

32 text to video models and apps on inference shell.run via API, SDK, or belt CLI.

32apps
1category

video

32
grok-imagine-video-1-5

grok-imagine-video-1-5

xai/grok-imagine-video-1-5

generate videos with xai's grok imagine video 1.5. text-to-video, image-to-video and reference-to-video at 480p, 720p or 1080p, 1-15 seconds, with generated audio.

video
MiniMax H3 Max

MiniMax H3 Max

falai/minimax-h3-max

minimax h3 max — fal's post-trained h3 for stronger prompt adherence and aesthetics. text-to-video, image-to-video with end frame, and reference-based generation. 480p/768p, 5-15s, ~3s per 5s clip.

video
Seedance 2.5

Seedance 2.5

bytedance/seedance-2-5

professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.5 model via byteplus ark api. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

video
seedance-2-5-studio

seedance-2-5-studio

bytedance/seedance-2-5-studio

professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

video
FLUX 3 Video

FLUX 3 Video

bfl/flux-3-video

flux 3 video by black forest labs — generate, animate, and extend video up to 20s at hd or full hd with synchronized audio. supports text-to-video, image-to-video with keyframes, video continuation, and draft mode.

video
MiniMax H3

MiniMax H3

minimax/h3

minimax h3 — multimodal video generation with native audio. text-to-video, image-to-video with end_image, and reference-based generation. 2k resolution, 5-15s duration, 24fps. prompt with timelines, audio cues, and negative lists.

video
Runway Gen-4.5

Runway Gen-4.5

runway/gen-4-5

runway gen-4.5 — high-quality video generation from text or image. 2-10 second duration, multiple aspect ratios. 12 credits/second.

video
PixVerse C1

PixVerse C1

pixverse/c1

pixverse c1 — advanced video generation model. text-to-video and image-to-video with up to 1080p resolution and 15s duration. per-second billing.

video
PixVerse V6

PixVerse V6

pixverse/v6

pixverse v6 — versatile video generation model. text-to-video and image-to-video with up to 1080p resolution and 15s duration. per-second billing.

video
seedance-2-0-studio-mini

seedance-2-0-studio-mini

bytedance/seedance-2-0-studio-mini

cost-effective multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Seedance 2.0 Mini

Seedance 2.0 Mini

bytedance/seedance-2-0-mini

cost-effective multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 mini model via byteplus ark api. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Gemini Omni Flash

Gemini Omni Flash

google/gemini-omni-flash

gemini omni flash — text-to-video with synchronized audio, grounded in real-world knowledge

video
HeyGen Video Agent

HeyGen Video Agent

heygen/video-agent

generate complete videos from natural language prompts using heygen's ai video agent. the agent handles avatar selection, scripting, and production automatically.

video
Kling Video V3

Kling Video V3

klingai/video-v3

kling v3.0 - latest and most capable video generation model. native 4k output, multi-shot generation, flexible 3-15s duration billed per second, element control, motion control, and synchronized audio.

video
Kling Video V2.6

Kling Video V2.6

klingai/video-v2-6

kling v2.6 video generation with native sound and voice control. supports text-to-video and image-to-video with start/end frames, synchronized audio generation, and voice-driven character animation.

video
Kling Video O1

Kling Video O1

klingai/video-o1

kling video o1 (omni) - unified video generation with text, image references, start/end frames, element references, and video references for editing and style transfer. the most capable kling model.

video
Kling Video V2.5 Turbo

Kling Video V2.5 Turbo

klingai/video-v2-5

kling v2.5 turbo - fast video generation from text and images. supports start/end frame interpolation in pro mode. optimized for speed while maintaining high quality at up to 1080p.

video
seedance-2-0-studio

seedance-2-0-studio

bytedance/seedance-2-0-studio

professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
seedance-2-0-studio-fast

seedance-2-0-studio-fast

bytedance/seedance-2-0-studio-fast

fast multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Seedance 2.0

Seedance 2.0

bytedance/seedance-2-0

professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 model via byteplus ark api. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Seedance 2.0 Fast

Seedance 2.0 Fast

bytedance/seedance-2-0-fast

fast multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 fast model via byteplus ark api. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
HappyHorse 1.0 Text-to-Video

HappyHorse 1.0 Text-to-Video

alibaba/happyhorse-1-0-t2v

happyhorse 1.0 text-to-video generates physically realistic videos with smooth motion from text prompts via dashscope api, supporting 720p/1080p resolution and up to 15 seconds duration

video
Wan 2.7 Text-to-Video

Wan 2.7 Text-to-Video

alibaba/wan-2-7-t2v

wan 2.7 text-to-video generates high-quality videos from text prompts using alibaba's latest video generation model via dashscope api, supporting 720p/1080p resolution and up to 15 seconds duration

video
Veo 3.1 Lite

Veo 3.1 Lite

google/veo-3-1-lite

veo 3.1 lite via gemini api - lightweight video generation with text and image input, audio support

video
Pruna Wan Text-to-Video

Pruna Wan Text-to-Video

pruna/wan-t2v

generate videos directly from text descriptions in 480p or 720p

video
P-Video

P-Video

pruna/p-video

fast text-to-video and image-to-video in 720p/1080p with audio support

video
Grok Imagine Video

Grok Imagine Video

xai/grok-imagine-video

generate and edit videos using xai's grok imagine video model. supports text-to-video, image-to-video, and video editing with configurable duration and resolution.

video
Veo 3.1

Veo 3.1

google/veo-3-1

veo 3.1 via vertex ai - advanced video generation with frame interpolation, reference images, and audio generation

video
Veo 3.1 Fast

Veo 3.1 Fast

google/veo-3-1-fast

veo 3.1 fast via vertex ai - generate videos from text prompts or images with optional audio

video
Seedance 1.0 Pro

Seedance 1.0 Pro

bytedance/seedance-1-0-pro

generate high-quality videos up to 1080p from text prompts with optional first-frame image control using bytedance's seedance 1.0 pro model.

video
Seedance 1.0 Pro Fast

Seedance 1.0 Pro Fast

bytedance/seedance-1-0-pro-fast

fast high-quality video generation up to 1080p from text prompts with optional first-frame image control using bytedance's seedance 1.0 pro fast model.

video
Seedance 1.5 Pro

Seedance 1.5 Pro

bytedance/seedance-1-5-pro

generate high-quality videos from text prompts with optional first-frame image control using bytedance's seedance 1.5 pro model via byteplus ark api.

video

explore more on inference shell

text to video is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.