video
generate and edit video, from text, images or other video. 88 apps on inference shell.
run via API, SDK, or belt CLI.

Seedance 2.5
Professional multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.5 model via BytePlus ARK API. Supports up to 4K (10-bit color), durations up to 30s, MOV output, and multimodal reference-to-video with synchronized audio.
video
P-Video-Avatar
Generate talking head videos from a portrait image with text or audio-driven speech
video
grok-imagine-video-1-5
xai/grok-imagine-video-1-5
generate videos with xai's grok imagine video 1.5. text-to-video, image-to-video and reference-to-video at 480p, 720p or 1080p, 1-15 seconds, with generated audio.

MiniMax H3 Max
falai/minimax-h3-max
minimax h3 max — fal's post-trained h3 for stronger prompt adherence and aesthetics. text-to-video, image-to-video with end frame, and reference-based generation. 480p/768p, 5-15s, ~3s per 5s clip.

Seedance 2.5
bytedance/seedance-2-5
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.5 model via byteplus ark api. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

seedance-2-5-studio
bytedance/seedance-2-5-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

FLUX 3 Video
bfl/flux-3-video
flux 3 video by black forest labs — generate, animate, and extend video up to 20s at hd or full hd with synchronized audio. supports text-to-video, image-to-video with keyframes, video continuation, and draft mode.

MiniMax H3
minimax/h3
minimax h3 — multimodal video generation with native audio. text-to-video, image-to-video with end_image, and reference-based generation. 2k resolution, 5-15s duration, 24fps. prompt with timelines, audio cues, and negative lists.

Runway Gen-4.5
runway/gen-4-5
runway gen-4.5 — high-quality video generation from text or image. 2-10 second duration, multiple aspect ratios. 12 credits/second.

PixVerse C1
pixverse/c1
pixverse c1 — advanced video generation model. text-to-video and image-to-video with up to 1080p resolution and 15s duration. per-second billing.

PixVerse V6
pixverse/v6
pixverse v6 — versatile video generation model. text-to-video and image-to-video with up to 1080p resolution and 15s duration. per-second billing.

seedance-2-0-studio-mini
bytedance/seedance-2-0-studio-mini
cost-effective multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Mini
bytedance/seedance-2-0-mini
cost-effective multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 mini model via byteplus ark api. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Gemini Omni Flash
google/gemini-omni-flash
gemini omni flash — text-to-video with synchronized audio, grounded in real-world knowledge

HeyGen Video Agent
heygen/video-agent
generate complete videos from natural language prompts using heygen's ai video agent. the agent handles avatar selection, scripting, and production automatically.

Kling Video V3
klingai/video-v3
kling v3.0 - latest and most capable video generation model. native 4k output, multi-shot generation, flexible 3-15s duration billed per second, element control, motion control, and synchronized audio.

Kling Video V2.6
klingai/video-v2-6
kling v2.6 video generation with native sound and voice control. supports text-to-video and image-to-video with start/end frames, synchronized audio generation, and voice-driven character animation.

Kling Video O1
klingai/video-o1
kling video o1 (omni) - unified video generation with text, image references, start/end frames, element references, and video references for editing and style transfer. the most capable kling model.

Kling Video V2.5 Turbo
klingai/video-v2-5
kling v2.5 turbo - fast video generation from text and images. supports start/end frame interpolation in pro mode. optimized for speed while maintaining high quality at up to 1080p.

seedance-2-0-studio
bytedance/seedance-2-0-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

seedance-2-0-studio-fast
bytedance/seedance-2-0-studio-fast
fast multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0
bytedance/seedance-2-0
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 model via byteplus ark api. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Fast
bytedance/seedance-2-0-fast
fast multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 fast model via byteplus ark api. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

HappyHorse 1.0 Text-to-Video
alibaba/happyhorse-1-0-t2v
happyhorse 1.0 text-to-video generates physically realistic videos with smooth motion from text prompts via dashscope api, supporting 720p/1080p resolution and up to 15 seconds duration

Wan 2.7 Text-to-Video
alibaba/wan-2-7-t2v
wan 2.7 text-to-video generates high-quality videos from text prompts using alibaba's latest video generation model via dashscope api, supporting 720p/1080p resolution and up to 15 seconds duration

Veo 3.1 Lite
google/veo-3-1-lite
veo 3.1 lite via gemini api - lightweight video generation with text and image input, audio support

Grok Reference Video
xai/grok-reference-video
generate videos using reference images for style and content guidance with xai's grok imagine video model. provide reference images to influence the visual style of generated videos.

Pruna Wan Text-to-Video
pruna/wan-t2v
generate videos directly from text descriptions in 480p or 720p

P-Video
pruna/p-video
fast text-to-video and image-to-video in 720p/1080p with audio support

Grok Imagine Video
xai/grok-imagine-video
generate and edit videos using xai's grok imagine video model. supports text-to-video, image-to-video, and video editing with configurable duration and resolution.

Veo 3.1
google/veo-3-1
veo 3.1 via vertex ai - advanced video generation with frame interpolation, reference images, and audio generation

Veo 3.1 Fast
google/veo-3-1-fast
veo 3.1 fast via vertex ai - generate videos from text prompts or images with optional audio

Seedance 1.0 Pro
bytedance/seedance-1-0-pro
generate high-quality videos up to 1080p from text prompts with optional first-frame image control using bytedance's seedance 1.0 pro model.

Seedance 1.0 Pro Fast
bytedance/seedance-1-0-pro-fast
fast high-quality video generation up to 1080p from text prompts with optional first-frame image control using bytedance's seedance 1.0 pro fast model.

Seedance 1.5 Pro
bytedance/seedance-1-5-pro
generate high-quality videos from text prompts with optional first-frame image control using bytedance's seedance 1.5 pro model via byteplus ark api.

Wan 2.5
falai/wan-2-5
creates high-quality, animated videos instantly from any static image.

P-Video-Edit
pruna/p-video-edit
edit videos from text prompts with optional reference images. up to 15s input, standard or draft quality.

Runway Act-Two
runway/act-two
runway act-two — character performance transfer. animate a character image or video using a reference performance video with facial expressions, gestures, and body control. 5 credits/second.

Runway Aleph 2.0
runway/aleph-2
runway aleph 2.0 — video-to-video transformation with text and keyframe guidance. supports aspect ratio targeting and up to 5 reference keyframes. 28 credits/second.

Mirage Text Overlays
mirage/text-overlays
render up to 4 static text variants onto one video with per-variant font, size and colour — built for testing ad hooks and headlines against the same footage.

PixVerse Modify
pixverse/modify
pixverse modify — edit video content using text prompts. swap subjects, add elements, or restyle regions with reference images.

P-Video-Replace
pruna/p-video-replace
replace characters in videos using reference images. preserves motion, timing, camera, and scene.

Bria Video Eraser
bria/video-eraser
erase objects from video using a mask with inpainting

Bria Video Replace Background
bria/video-replace-background
replace video background with an image or another video

Bria Video Green Screen
bria/video-green-screen
apply green or blue screen effect to video foreground

HappyHorse 1.0 Video Edit
alibaba/happyhorse-1-0-video-edit
happyhorse 1.0 video edit supports advanced video editing through natural language instructions with up to 5 reference images, preserving original motion dynamics via dashscope api

Wan 2.7 Video Edit
alibaba/wan-2-7-videoedit
wan 2.7 video edit performs instruction-based video editing and style transfer using multimodal inputs (text, images, video) via dashscope api with 720p/1080p output

Mirage Avatar X
mirage/avatar-x
generate an expressive talking-head video from a portrait image or video reference and an audio track using mirage avatar x, with natural lip sync, eye contact and micro-expressions.

Mirage Video 1
mirage/video-1
generate an expressive talking-head video from a portrait image and an audio track using mirage video 1, with natural lip sync, eye contact and micro-expressions.

PixVerse Lip Sync
pixverse/lip-sync
pixverse lip sync — align speech to mouth movement in video. supports audio file input or text-to-speech with selectable voices.

HeyGen Lip Sync
heygen/lipsync
re-sync video lip movements to new audio using heygen's lipsync technology. supports speed and precision modes with optional captioning.

HeyGen Video Translate
heygen/video-translate
translate videos into 30+ languages with voice cloning and lip-sync using heygen. supports speed and precision modes with optional captioning.

Kling Lip Sync
klingai/lip-sync
kling lip sync - drive mouth movements in videos using text or audio. ideal for dubbing, adding speech to silent videos, or replacing dialogue.

OmniHuman 1.5
bytedance/omnihuman-1-5
multi-character audio-driven avatar video generation. takes a portrait image + audio and generates a video where the person speaks/sings in sync. supports specifying which character to drive.

OmniHuman 1.0
bytedance/omnihuman-1-0
audio-driven avatar video generation. takes a portrait image + audio and generates a video where the person speaks/sings in sync with the audio.

Fabric 1.0
falai/fabric-1-0
creates videos where an image appears to talk using advanced lip-sync technology.

Runway Gen-4 Turbo
runway/gen-4-turbo
runway gen-4 turbo — fast image-to-video generation. animate images into video with text guidance. 5 credits/second.

HeyGen Photo Video
heygen/photo-video
animate portrait photos into talking videos using heygen. upload a face image and add speech with configurable voice, motion prompts, and expressiveness.

HappyHorse 1.0 Image-to-Video
alibaba/happyhorse-1-0-i2v
happyhorse 1.0 image-to-video generates physically realistic videos with smooth motion from a single image and optional text description via dashscope api, supporting 720p/1080p resolution

Wan 2.7 Image-to-Video
alibaba/wan-2-7-i2v
wan 2.7 image-to-video generates videos from images using multi-modal input (text, images, audio, video). supports first frame generation, first+last frame, and video continuation with 720p/1080p resolution

Pruna Wan Image-to-Video
pruna/wan-i2v
transform static images into animated videos with text prompts

Wan 2.5 Image-to-Video
falai/wan-2-5-i2v
generates high-quality video content from static images.

PixVerse Avatar
pixverse/avatar
pixverse avatar — generate talking avatar video from a portrait image. supports audio file or text-to-speech input at up to 1080p.

HeyGen Create Avatar
heygen/create-avatar
create heygen avatars from video footage (digital twin), a photo (photo avatar), or a text prompt (ai-generated). returns a look id for use with avatar-video.

HeyGen Avatar Video
heygen/avatar-video
generate talking avatar videos using heygen's digital and photo avatars with avatar iv or v engines, configurable voice, resolution up to 4k, and expressiveness.

Kling Avatar
klingai/avatar
kling avatar - generate digital human broadcast-style talking head videos from a single face photo. provide text or audio for the avatar to speak.

P-Video-Avatar
pruna/p-video-avatar
generate talking head videos from a portrait image with text or audio-driven speech

Topaz Astra
topaz/astra
creative video upscaling — ai-guided upscaling with prompt and creativity controls

Topaz Starlight
topaz/starlight
generative video upscaling — precision, hq, mini, sharp, and fast models

Topaz Proteus
topaz/proteus
proteus video upscaling and enhancement — precision upscaling, deinterlacing, face recovery, and cgi enhancement

Bria Video Increase Resolution
bria/video-increase-resolution
upscale video resolution up to 8k using ai super-resolution

PixVerse Fusion
pixverse/fusion
pixverse fusion — generate video from reference images. compose 1-3 subject/background images into a video scene with text-guided motion.

HappyHorse 1.0 Reference-to-Video
alibaba/happyhorse-1-0-r2v
happyhorse 1.0 reference-to-video generates videos preserving subject characters from up to 9 reference images, with enhanced stability in subject and scene referencing via dashscope api

Wan 2.7 Reference-to-Video
alibaba/wan-2-7-r2v
wan 2.7 reference-to-video generates videos featuring characters from reference images and videos, supporting multi-character interaction, voice timbre cloning, and first-frame control

HTML to Video
infsh/html-to-video
render html/css/js animations to video — supports gsap timelines, css animations, web animations api

HeyGen Hyperframes Render
infsh/hyperframes-render
render heygen hyperframes compositions to video — supports clips, gsap timelines, track layering

Remotion Render
infsh/remotion-render
render videos from react/remotion component code — pass tsx, get mp4

VideoSeal Watermarking
infsh/videoseal
embed or detect imperceptible watermarking for images and videos.

Inference Shell Extract Last Frame
infsh/extract-last-frame
save a specific frame from the end of a video as a static image file.

Inference Shell Media Merger
infsh/media-merger
merge, concatenate, and stitch multiple videos and images into a single video with customizable transitions. supports common formats (mp4, mov, webm, png, jpg), custom frame rates, and transition effects between clips. built on ffmpeg

Mirage Video Captions
mirage/video-captions
burn styled animated captions onto a vertical video using mirage caption templates, from an upload or an existing mirage video id.

VEED Subtitles
veed/subtitles
add professional burned-in subtitles to videos with 25+ style presets. supports 100+ languages with automatic transcription or custom srt files.

PixVerse Extend
pixverse/extend
pixverse extend — extend an existing video with ai-generated continuation. supports 5s or 8s extensions at up to 1080p.

Grok Video Extender
xai/grok-extend-video
extend existing videos using xai's grok imagine video model. takes an existing video and generates additional frames to continue it with prompt guidance.

Topaz Frame Interpolation
topaz/frame-interpolation
video frame interpolation — slowmo and fps boost with apollo, chronos, and aion models

RIFE Video Interpolation
infsh/rife-video-interpolation
increases video smoothness by using ai to generate and insert extra frames, effectively raising the video's frame rate.
explore more on inference shell
video is one of the categories on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.



