text to speech
22 text to speech models and apps on inference shell.
run via API, SDK, or belt CLI.
audio
22
Eleven Flash v2.5
elevenlabs/eleven-flash-v2-5
elevenlabs eleven flash v2.5 text to speech. the fastest and lowest-cost elevenlabs model, for long texts and high volume.

Eleven Multilingual v2
elevenlabs/eleven-multilingual-v2
elevenlabs eleven multilingual v2 text to speech. stable, high-quality multilingual speech for long-form content.

Eleven v3
elevenlabs/eleven-v3
elevenlabs eleven v3 text to speech. expressive speech in 70+ languages with audio tags for emotion and delivery.

Eleven v4 Turbo
elevenlabs/eleven-v4-turbo
elevenlabs eleven v4 turbo text to speech. the low-latency v4 model; send a text for an audio file, or stream text over a socket and hear it as it is written.

Eleven v4
elevenlabs/eleven-v4
elevenlabs eleven v4 text to speech. the newest and most expressive model, with 90+ languages and audio tags; send a text for an audio file, or stream text over a socket and hear it as it is written.

HeyGen Instant Voice Clone
heygen/voice-clone
clone a voice from a single short recording with heygen instant clone, and speak any text in it in the same call. the cloned voice id also works with heygen/text-to-speech and avatar videos.

DramaBox Text-to-Speech
infsh/dramabox
expressive text-to-speech with voice cloning. the prompt controls emotion, laughs, sighs, pauses and delivery.

MiniMax Speech 2.8 HD
minimax/speech-2-8-hd
minimax speech 2.8 hd — high-quality text-to-speech. 40 languages, 9 emotions, custom voices. direct minimax api.

MiniMax Speech 2.8 Turbo
minimax/speech-2-8-turbo
minimax speech 2.8 turbo — fast text-to-speech. 40 languages, 9 emotions, custom voices. lower latency than hd. direct minimax api.

HeyGen Text to Speech
heygen/text-to-speech
generate natural speech audio from text using heygen's starfish tts engine. supports configurable voice, speed, ssml input, and multiple languages.

Inworld TTS-2
inworld/text-to-speech-2
inworld tts-2 - high-quality multilingual text-to-speech with 100+ languages and natural-language steering

Inworld TTS 1.5 Max
inworld/text-to-speech-1-5-max
inworld tts 1.5 max - low-latency text-to-speech with 15 languages (<200ms p50)

Inworld TTS 1.5 Mini
inworld/text-to-speech-1-5-mini
inworld tts 1.5 mini - ultra-low-latency text-to-speech with 15 languages (~120ms p50)

OmniVoice TTS
infsh/omnivoice
zero-shot text-to-speech with voice cloning for 600+ languages.

Grok TTS
xai/grok-tts
convert text into natural speech using xai's text to speech api. supports multiple voices, expressive speech tags, and mp3/wav/pcm output formats.

ElevenLabs Text to Dialogue
elevenlabs/text-to-dialogue
elevenlabs text to dialogue - generate immersive multi-voice dialogue

ElevenLabs Text to Speech
elevenlabs/tts
elevenlabs text to speech - high-quality multilingual voice synthesis

Kokoro TTS
falai/kokoro-tts
kokoro tts - lightweight text-to-speech with multiple languages and voices

Dia TTS
falai/dia-tts
dia tts - generate realistic dialogue with emotion control, natural nonverbals, and voice cloning

Dia TTS
infsh/dia-tts
dia tts - generate realistic dialogue with emotion control, natural nonverbals, and voice cloning

Higgs Audio
infsh/higgs-audio
generates speech from text with advanced, expressive audio quality.

Chatterbox TTS
infsh/chatterbox
converts written text into spoken audio.
explore more on inference shell
text to speech is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.