native audio
17 native audio models and apps on inference shell.
run via API, SDK, or belt CLI.

Seedance 2.5
Professional multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.5 model via BytePlus ARK API. Supports up to 4K (10-bit color), durations up to 30s, MOV output, and multimodal reference-to-video with synchronized audio.
video
Veo 3.1 Fast
Veo 3.1 Fast via Vertex AI - Generate videos from text prompts or images with optional audio
videovideo
17
Seedance 2.5
bytedance/seedance-2-5
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.5 model via byteplus ark api. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

seedance-2-5-studio
bytedance/seedance-2-5-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

FLUX 3 Video
bfl/flux-3-video
flux 3 video by black forest labs — generate, animate, and extend video up to 20s at hd or full hd with synchronized audio. supports text-to-video, image-to-video with keyframes, video continuation, and draft mode.

MiniMax H3
minimax/h3
minimax h3 — multimodal video generation with native audio. text-to-video, image-to-video with end_image, and reference-based generation. 2k resolution, 5-15s duration, 24fps. prompt with timelines, audio cues, and negative lists.

seedance-2-0-studio-mini
bytedance/seedance-2-0-studio-mini
cost-effective multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Mini
bytedance/seedance-2-0-mini
cost-effective multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 mini model via byteplus ark api. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Gemini Omni Flash
google/gemini-omni-flash
gemini omni flash — text-to-video with synchronized audio, grounded in real-world knowledge

Kling Video V3
klingai/video-v3
kling v3.0 - latest and most capable video generation model. native 4k output, multi-shot generation, flexible 3-15s duration billed per second, element control, motion control, and synchronized audio.

Kling Video V2.6
klingai/video-v2-6
kling v2.6 video generation with native sound and voice control. supports text-to-video and image-to-video with start/end frames, synchronized audio generation, and voice-driven character animation.

seedance-2-0-studio
bytedance/seedance-2-0-studio
professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

seedance-2-0-studio-fast
bytedance/seedance-2-0-studio-fast
fast multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0
bytedance/seedance-2-0
professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 model via byteplus ark api. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Seedance 2.0 Fast
bytedance/seedance-2-0-fast
fast multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 fast model via byteplus ark api. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

Veo 3.1 Lite
google/veo-3-1-lite
veo 3.1 lite via gemini api - lightweight video generation with text and image input, audio support

P-Video
pruna/p-video
fast text-to-video and image-to-video in 720p/1080p with audio support

Veo 3.1
google/veo-3-1
veo 3.1 via vertex ai - advanced video generation with frame interpolation, reference images, and audio generation

Veo 3.1 Fast
google/veo-3-1-fast
veo 3.1 fast via vertex ai - generate videos from text prompts or images with optional audio
explore more on inference shell
native audio is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.