# audio apps

> speech, music, sound effects and transcription. 44 apps on inference shell.

**URL:** https://inference.sh/apps/category/audio

---

- [Eleven Flash v2.5](https://inference.sh/apps/elevenlabs/eleven-flash-v2-5.md) `elevenlabs/eleven-flash-v2-5`: $0.04/1K characters
- [Eleven Multilingual v2](https://inference.sh/apps/elevenlabs/eleven-multilingual-v2.md) `elevenlabs/eleven-multilingual-v2`: $0.08/1K characters
- [Eleven v3](https://inference.sh/apps/elevenlabs/eleven-v3.md) `elevenlabs/eleven-v3`: $0.08/1K characters
- [Eleven v4 Turbo](https://inference.sh/apps/elevenlabs/eleven-v4-turbo.md) `elevenlabs/eleven-v4-turbo`: $0.04/1K characters
- [Eleven v4](https://inference.sh/apps/elevenlabs/eleven-v4.md) `elevenlabs/eleven-v4`: $0.08/1K characters
- [GPT Transcribe](https://inference.sh/apps/openai/gpt-transcribe.md) `openai/gpt-transcribe`: $0.0045/min of audio for files, $0.017/min for realtime
- [Grok Speech to Text](https://inference.sh/apps/xai/grok-stt.md) `xai/grok-stt`: $0.10/hr of audio for files, $0.20/hr for realtime
- [grok-voice](https://inference.sh/apps/xai/grok-voice.md) `xai/grok-voice`: $0.08/min session audio + $0.004/text input
- [HeyGen Instant Voice Clone](https://inference.sh/apps/heygen/voice-clone.md) `heygen/voice-clone`: $0.0007/sec audio (clone-only runs: free)
- [DramaBox Text-to-Speech](https://inference.sh/apps/infsh/dramabox.md) `infsh/dramabox`: GPU time, billed per millisecond
- [MiniMax Music Cover](https://inference.sh/apps/minimax/music-cover.md) `minimax/music-cover`: $0.15/cover
- [MiniMax Speech 2.8 HD](https://inference.sh/apps/minimax/speech-2-8-hd.md) `minimax/speech-2-8-hd`: $0.0001/char
- [MiniMax Music 3.0](https://inference.sh/apps/minimax/music-3-0.md) `minimax/music-3-0`: $0.15/track
- [MiniMax Speech 2.8 Turbo](https://inference.sh/apps/minimax/speech-2-8-turbo.md) `minimax/speech-2-8-turbo`: $60.00/M chars
- [ElevenLabs Voice Remix](https://inference.sh/apps/elevenlabs/voice-remix.md) `elevenlabs/voice-remix`: $0.10/remix
- [ElevenLabs Voice Clone](https://inference.sh/apps/elevenlabs/voice-clone.md) `elevenlabs/voice-clone`: $0.10/clone
- [ElevenLabs Voice Design](https://inference.sh/apps/elevenlabs/voice-design.md) `elevenlabs/voice-design`: $0.10/design
- [HeyGen Text to Speech](https://inference.sh/apps/heygen/text-to-speech.md) `heygen/text-to-speech`: $0.0007/sec
- [Kling Video to Audio](https://inference.sh/apps/klingai/video-to-audio.md) `klingai/video-to-audio`: $0.0088/video
- [Inworld Voice Cloning](https://inference.sh/apps/inworld/voice-cloning.md) `inworld/voice-cloning`: from $0.000003/sec
- [Inworld Voice Design](https://inference.sh/apps/inworld/voice-design.md) `inworld/voice-design`: from $0.000003/sec
- [Inworld TTS-2](https://inference.sh/apps/inworld/text-to-speech-2.md) `inworld/text-to-speech-2`: $35.00/M chars
- [Inworld TTS 1.5 Max](https://inference.sh/apps/inworld/text-to-speech-1-5-max.md) `inworld/text-to-speech-1-5-max`: $35.00/M chars
- [Inworld Speech to Text](https://inference.sh/apps/inworld/speech-to-text.md) `inworld/speech-to-text`: $0.35/hr of audio
- [Inworld TTS 1.5 Mini](https://inference.sh/apps/inworld/text-to-speech-1-5-mini.md) `inworld/text-to-speech-1-5-mini`: $25.00/M chars
- [OmniVoice TTS](https://inference.sh/apps/infsh/omnivoice.md) `infsh/omnivoice`: GPU time, billed per millisecond
- [Grok TTS](https://inference.sh/apps/xai/grok-tts.md) `xai/grok-tts`: $15.00/M chars
- [ElevenLabs Forced Alignment](https://inference.sh/apps/elevenlabs/forced-alignment.md) `elevenlabs/forced-alignment`: $0.48/hr audio
- [ElevenLabs Text to Dialogue](https://inference.sh/apps/elevenlabs/text-to-dialogue.md) `elevenlabs/text-to-dialogue`: $0.05/1K chars (Flash/Turbo), $0.10/1K chars (Multilingual v2/v3)
- [ElevenLabs Dubbing](https://inference.sh/apps/elevenlabs/dubbing.md) `elevenlabs/dubbing`: $0.33/min (watermark), $0.50/min (no watermark)
- [ElevenLabs Music](https://inference.sh/apps/elevenlabs/music.md) `elevenlabs/music`: $0.15/min
- [ElevenLabs Sound Effects](https://inference.sh/apps/elevenlabs/sound-effects.md) `elevenlabs/sound-effects`: $0.12/min
- [ElevenLabs Voice Isolator](https://inference.sh/apps/elevenlabs/voice-isolator.md) `elevenlabs/voice-isolator`: $0.12/min
- [ElevenLabs Voice Changer](https://inference.sh/apps/elevenlabs/voice-changer.md) `elevenlabs/voice-changer`: $0.12/min
- [ElevenLabs Speech to Text](https://inference.sh/apps/elevenlabs/stt.md) `elevenlabs/stt`: $0.22-$0.39/hr (Scribe v2: $0.22/hr, Realtime: $0.39/hr)
- [ElevenLabs Text to Speech](https://inference.sh/apps/elevenlabs/tts.md) `elevenlabs/tts`: $0.10/1K chars (v2), $0.08/1K chars (v4), $0.04/1K chars (v4 Turbo), $0.05/1K chars (Flash/Turbo)
- [Kokoro TTS](https://inference.sh/apps/falai/kokoro-tts.md) `falai/kokoro-tts`: $0.02/1K chars
- [Dia TTS](https://inference.sh/apps/falai/dia-tts.md) `falai/dia-tts`: $0.04/K chars
- [Dia TTS](https://inference.sh/apps/infsh/dia-tts.md) `infsh/dia-tts`: from $0.000003/sec
- [DiffRhythm Song Generator](https://inference.sh/apps/infsh/diffrythm.md) `infsh/diffrythm`: GPU time, billed per millisecond
- [Audio-X Generator](https://inference.sh/apps/infsh/audio-x.md) `infsh/audio-x`: GPU time, billed per millisecond
- [Kokoro TTS](https://inference.sh/apps/infsh/kokoro-tts.md) `infsh/kokoro-tts`: from $0.001/sec
- [Higgs Audio](https://inference.sh/apps/infsh/higgs-audio.md) `infsh/higgs-audio`: GPU time, billed per millisecond
- [Chatterbox TTS](https://inference.sh/apps/infsh/chatterbox.md) `infsh/chatterbox`: GPU time, billed per millisecond