apps/native audio

native audio

17 native audio models and apps on inference shell.run via API, SDK, or belt CLI.

17apps
1category

video

17
Seedance 2.5

Seedance 2.5

bytedance/seedance-2-5

professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.5 model via byteplus ark api. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

video
seedance-2-5-studio

seedance-2-5-studio

bytedance/seedance-2-5-studio

professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), durations up to 30s, mov output, and multimodal reference-to-video with synchronized audio.

video
FLUX 3 Video

FLUX 3 Video

bfl/flux-3-video

flux 3 video by black forest labs — generate, animate, and extend video up to 20s at hd or full hd with synchronized audio. supports text-to-video, image-to-video with keyframes, video continuation, and draft mode.

video
MiniMax H3

MiniMax H3

minimax/h3

minimax h3 — multimodal video generation with native audio. text-to-video, image-to-video with end_image, and reference-based generation. 2k resolution, 5-15s duration, 24fps. prompt with timelines, audio cues, and negative lists.

video
seedance-2-0-studio-mini

seedance-2-0-studio-mini

bytedance/seedance-2-0-studio-mini

cost-effective multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Seedance 2.0 Mini

Seedance 2.0 Mini

bytedance/seedance-2-0-mini

cost-effective multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 mini model via byteplus ark api. ~50% cheaper than seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Gemini Omni Flash

Gemini Omni Flash

google/gemini-omni-flash

gemini omni flash — text-to-video with synchronized audio, grounded in real-world knowledge

video
Kling Video V3

Kling Video V3

klingai/video-v3

kling v3.0 - latest and most capable video generation model. native 4k output, multi-shot generation, flexible 3-15s duration billed per second, element control, motion control, and synchronized audio.

video
Kling Video V2.6

Kling Video V2.6

klingai/video-v2-6

kling v2.6 video generation with native sound and voice control. supports text-to-video and image-to-video with start/end frames, synchronized audio generation, and voice-driven character animation.

video
seedance-2-0-studio

seedance-2-0-studio

bytedance/seedance-2-0-studio

professional multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
seedance-2-0-studio-fast

seedance-2-0-studio-fast

bytedance/seedance-2-0-studio-fast

fast multimodal video generation with private asset library support. automatically uploads reference images to byteplus virtual portrait library for enhanced character consistency. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Seedance 2.0

Seedance 2.0

bytedance/seedance-2-0

professional multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 model via byteplus ark api. supports up to 4k (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Seedance 2.0 Fast

Seedance 2.0 Fast

bytedance/seedance-2-0-fast

fast multimodal video generation from text, images, video, and audio references using bytedance's seedance 2.0 fast model via byteplus ark api. supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio.

video
Veo 3.1 Lite

Veo 3.1 Lite

google/veo-3-1-lite

veo 3.1 lite via gemini api - lightweight video generation with text and image input, audio support

video
P-Video

P-Video

pruna/p-video

fast text-to-video and image-to-video in 720p/1080p with audio support

video
Veo 3.1

Veo 3.1

google/veo-3-1

veo 3.1 via vertex ai - advanced video generation with frame interpolation, reference images, and audio generation

video
Veo 3.1 Fast

Veo 3.1 Fast

google/veo-3-1-fast

veo 3.1 fast via vertex ai - generate videos from text prompts or images with optional audio

video

explore more on inference shell

native audio is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.