text to video
32 text to video models and apps on inference shell.
run via API, SDK, or belt CLI.

Seedance 2.5
Professional multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.5 model via BytePlus ARK API. Supports up to 4K (10-bit color), durations up to 30s, MOV output, and multimodal reference-to-video with synchronized audio.
video
Veo 3.1 Fast
Veo 3.1 Fast via Vertex AI - Generate videos from text prompts or images with optional audio
videovideo
32
grok-imagine-video-1-5
xai/grok-imagine-video-1-5
generate videos with xai's grok imagine video 1.5. text-to-video, image-to-video and reference-to-video at 480p, 720p or 1080p, 1-15 seconds, with generated audio.

MiniMax H3 Max
falai/minimax-h3-max
minimax h3 max — fal's post-trained h3 for stronger prompt adherence and aesthetics. text-to-video, image-to-video with end frame, and reference-based generation. 480p/768p, 5-15s, ~3s per 5s clip.

Seedance 2.5
bytedance/seedance-2-5
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.5 model via byteplus ark api. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

seedance-2-5-studio
bytedance/seedance-2-5-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

FLUX 3 Video
bfl/flux-3-video
flux 3 video by black forest labs — generate, animate, and extend video up to 20s at hd or full hd with synchronized audio. supports text-to-video, image-to-video with keyframes, video continuation, and draft mode.

MiniMax H3
minimax/h3
minimax h3 — multimodal video generation with native audio. text-to-video, image-to-video with end_image, and reference-based generation. 2k resolution, 5-15s duration, 24fps. prompt with timelines, audio cues, and negative lists.

Runway Gen-4.5
runway/gen-4-5
runway gen-4.5 — high-quality video generation from text or image. 2-10 second duration, multiple aspect ratios. 12 credits/second.

PixVerse C1
pixverse/c1
pixverse c1 — advanced video generation model. text-to-video and image-to-video with up to 1080p resolution and 15s duration. per-second billing.

PixVerse V6
pixverse/v6
pixverse v6 — versatile video generation model. text-to-video and image-to-video with up to 1080p resolution and 15s duration. per-second billing.

seedance-2-0-studio-mini
bytedance/seedance-2-0-studio-mini
cost-effective multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Mini
bytedance/seedance-2-0-mini
cost-effective multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 mini model via byteplus ark api. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Gemini Omni Flash
google/gemini-omni-flash
gemini omni flash — text-to-video with synchronized audio, grounded in real-world knowledge

HeyGen Video Agent
heygen/video-agent
generate complete videos from natural language prompts using heygen's ai video agent. the agent handles avatar selection, scripting, and production automatically.

Kling Video V3
klingai/video-v3
kling v3.0 - latest and most capable video generation model. native 4k output, multi-shot generation, flexible 3-15s duration billed per second, element control, motion control, and synchronized audio.

Kling Video V2.6
klingai/video-v2-6
kling v2.6 video generation with native sound and voice control. supports text-to-video and image-to-video with start/end frames, synchronized audio generation, and voice-driven character animation.

Kling Video O1
klingai/video-o1
kling video o1 (omni) - unified video generation with text, image references, start/end frames, element references, and video references for editing and style transfer. the most capable kling model.

Kling Video V2.5 Turbo
klingai/video-v2-5
kling v2.5 turbo - fast video generation from text and images. supports start/end frame interpolation in pro mode. optimized for speed while maintaining high quality at up to 1080p.

seedance-2-0-studio
bytedance/seedance-2-0-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

seedance-2-0-studio-fast
bytedance/seedance-2-0-studio-fast
fast multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0
bytedance/seedance-2-0
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 model via byteplus ark api. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Fast
bytedance/seedance-2-0-fast
fast multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 fast model via byteplus ark api. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

HappyHorse 1.0 Text-to-Video
alibaba/happyhorse-1-0-t2v
happyhorse 1.0 text-to-video generates physically realistic videos with smooth motion from text prompts via dashscope api, supporting 720p/1080p resolution and up to 15 seconds duration

Wan 2.7 Text-to-Video
alibaba/wan-2-7-t2v
wan 2.7 text-to-video generates high-quality videos from text prompts using alibaba's latest video generation model via dashscope api, supporting 720p/1080p resolution and up to 15 seconds duration

Veo 3.1 Lite
google/veo-3-1-lite
veo 3.1 lite via gemini api - lightweight video generation with text and image input, audio support

Pruna Wan Text-to-Video
pruna/wan-t2v
generate videos directly from text descriptions in 480p or 720p

P-Video
pruna/p-video
fast text-to-video and image-to-video in 720p/1080p with audio support

Grok Imagine Video
xai/grok-imagine-video
generate and edit videos using xai's grok imagine video model. supports text-to-video, image-to-video, and video editing with configurable duration and resolution.

Veo 3.1
google/veo-3-1
veo 3.1 via vertex ai - advanced video generation with frame interpolation, reference images, and audio generation

Veo 3.1 Fast
google/veo-3-1-fast
veo 3.1 fast via vertex ai - generate videos from text prompts or images with optional audio

Seedance 1.0 Pro
bytedance/seedance-1-0-pro
generate high-quality videos up to 1080p from text prompts with optional first-frame image control using bytedance's seedance 1.0 pro model.

Seedance 1.0 Pro Fast
bytedance/seedance-1-0-pro-fast
fast high-quality video generation up to 1080p from text prompts with optional first-frame image control using bytedance's seedance 1.0 pro fast model.

Seedance 1.5 Pro
bytedance/seedance-1-5-pro
generate high-quality videos from text prompts with optional first-frame image control using bytedance's seedance 1.5 pro model via byteplus ark api.
explore more on inference shell
text to video is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.