audio
speech, music, sound effects and transcription. 44 apps on inference shell.
run via API, SDK, or belt CLI.

Eleven Flash v2.5
elevenlabs/eleven-flash-v2-5
elevenlabs eleven flash v2.5 text to speech. the fastest and lowest-cost elevenlabs model, for long texts and high volume.

Eleven Multilingual v2
elevenlabs/eleven-multilingual-v2
elevenlabs eleven multilingual v2 text to speech. stable, high-quality multilingual speech for long-form content.

Eleven v3
elevenlabs/eleven-v3
elevenlabs eleven v3 text to speech. expressive speech in 70+ languages with audio tags for emotion and delivery.

Eleven v4 Turbo
elevenlabs/eleven-v4-turbo
elevenlabs eleven v4 turbo text to speech. the low-latency v4 model; send a text for an audio file, or stream text over a socket and hear it as it is written.

Eleven v4
elevenlabs/eleven-v4
elevenlabs eleven v4 text to speech. the newest and most expressive model, with 90+ languages and audio tags; send a text for an audio file, or stream text over a socket and hear it as it is written.

HeyGen Instant Voice Clone
heygen/voice-clone
clone a voice from a single short recording with heygen instant clone, and speak any text in it in the same call. the cloned voice id also works with heygen/text-to-speech and avatar videos.

DramaBox Text-to-Speech
infsh/dramabox
expressive text-to-speech with voice cloning. the prompt controls emotion, laughs, sighs, pauses and delivery.

MiniMax Speech 2.8 HD
minimax/speech-2-8-hd
minimax speech 2.8 hd — high-quality text-to-speech. 40 languages, 9 emotions, custom voices. direct minimax api.

MiniMax Speech 2.8 Turbo
minimax/speech-2-8-turbo
minimax speech 2.8 turbo — fast text-to-speech. 40 languages, 9 emotions, custom voices. lower latency than hd. direct minimax api.

HeyGen Text to Speech
heygen/text-to-speech
generate natural speech audio from text using heygen's starfish tts engine. supports configurable voice, speed, ssml input, and multiple languages.

Inworld TTS-2
inworld/text-to-speech-2
inworld tts-2 - high-quality multilingual text-to-speech with 100+ languages and natural-language steering

Inworld TTS 1.5 Max
inworld/text-to-speech-1-5-max
inworld tts 1.5 max - low-latency text-to-speech with 15 languages (<200ms p50)

Inworld TTS 1.5 Mini
inworld/text-to-speech-1-5-mini
inworld tts 1.5 mini - ultra-low-latency text-to-speech with 15 languages (~120ms p50)

OmniVoice TTS
infsh/omnivoice
zero-shot text-to-speech with voice cloning for 600+ languages.

Grok TTS
xai/grok-tts
convert text into natural speech using xai's text to speech api. supports multiple voices, expressive speech tags, and mp3/wav/pcm output formats.

ElevenLabs Text to Dialogue
elevenlabs/text-to-dialogue
elevenlabs text to dialogue - generate immersive multi-voice dialogue

ElevenLabs Text to Speech
elevenlabs/tts
elevenlabs text to speech - high-quality multilingual voice synthesis

Kokoro TTS
falai/kokoro-tts
kokoro tts - lightweight text-to-speech with multiple languages and voices

Dia TTS
falai/dia-tts
dia tts - generate realistic dialogue with emotion control, natural nonverbals, and voice cloning

Dia TTS
infsh/dia-tts
dia tts - generate realistic dialogue with emotion control, natural nonverbals, and voice cloning

Higgs Audio
infsh/higgs-audio
generates speech from text with advanced, expressive audio quality.

Chatterbox TTS
infsh/chatterbox
converts written text into spoken audio.

GPT Transcribe
openai/gpt-transcribe
transcribe speech with openai. send a recording and get its transcript from gpt transcribe, or stream a microphone over a socket and read the transcript as it is spoken with gpt live transcribe.

Grok Speech to Text
xai/grok-stt
transcribe speech with xai's grok speech to text. send a recording and get the transcript with word timings and speakers, or stream a microphone over a socket and read the transcript as it is spoken.

Inworld Speech to Text
inworld/speech-to-text
inworld speech to text - multi-provider speech transcription with word timestamps

ElevenLabs Speech to Text
elevenlabs/stt
elevenlabs speech to text (scribe) - high-accuracy transcription with diarization

MiniMax Music Cover
minimax/music-cover
minimax music cover — ai-powered song covers and style transfer. upload a reference track and generate a cover with new style and optional new lyrics.

MiniMax Music 3.0
minimax/music-3-0
minimax music 3.0 — ai music generation up to 5 minutes. supports vocals with lyrics, instrumentals, and style prompts. direct minimax api.

ElevenLabs Music
elevenlabs/music
elevenlabs music - generate studio-quality music from text prompts

DiffRhythm Song Generator
infsh/diffrythm
generates complete songs quickly and simply using advanced latent diffusion technology.

ElevenLabs Voice Remix
elevenlabs/voice-remix
elevenlabs voice remix - modify voice characteristics like accent, gender, style, pacing

ElevenLabs Voice Design
elevenlabs/voice-design
elevenlabs voice design - create custom ai voices from text descriptions

Inworld Voice Design
inworld/voice-design
design a custom voice from a text description using inworld ai. describe the voice you want and get up to 3 previews. publish the one you like to use with any inworld tts model.

Kling Video to Audio
klingai/video-to-audio
kling video-to-audio - add generated sound effects, ambient audio, or music to any video. works with kling-generated and user-uploaded videos (3-20s).

ElevenLabs Sound Effects
elevenlabs/sound-effects
elevenlabs sound effects - generate custom sound effects from text

Audio-X Generator
infsh/audio-x
generates audio from any input using a unified framework.

grok-voice
xai/grok-voice
talk with grok. stream microphone audio over a socket and hear xai's speech-to-speech model answer as it speaks, with both sides transcribed; the conversation is the task's output.

ElevenLabs Voice Changer
elevenlabs/voice-changer
elevenlabs voice changer - transform voice in audio to a different voice
explore more on inference shell
audio is one of the categories on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.





