apps/heygen/voice-clone

voice-clone

Clone a voice from a single short recording with HeyGen Instant Clone, and speak any text in it in the same call. The cloned voice ID also works with heygen/text-to-speech and avatar videos.

run with your agent
# install belt
$curl -fsSL https://cli.inference.sh | sh
# view schema & details
$belt app get heygen/voice-clone
# run
$belt app run heygen/voice-clone

api reference

about

clone a voice from a single short recording with heygen instant clone, and speak any text in it in the same call. the cloned voice id also works with heygen/text-to-speech and avatar videos.

1. calling the api

install the client

the client provides a convenient way to interact with the api.

bash
1pip install inferencesh

setup your api key

set INFERENCE_API_KEY as an environment variable. get your key from settings → api keys.

bash
1export INFERENCE_API_KEY="inf_your_key"

run and get result

submit a request and wait for the final result. best for batch processing or when you don't need progress updates.

python
1from inferencesh import inference23client = inference()456result = client.run({7        "app": "heygen/voice-clone",8        "input": {}9    })1011print(result["output"])

stream live updates

get real-time progress updates as the task runs. ideal for showing progress bars, partial results, or long-running tasks.

python
1from inferencesh import inference23client = inference()456# stream=True yields updates as they arrive7for update in client.run({8        "app": "heygen/voice-clone",9        "input": {}10    }, stream=True):11    if update.get("progress"):12        print(f"progress: {update['progress']}%")13    if update.get("output"):14        print(f"output: {update['output']}")

2. authentication

the api uses api keys for authentication. see the authentication docs for detailed setup instructions.

3. files

file inputs are automatically handled by the sdk. you can pass local paths, urls, or base64 data.

automatic upload

the python sdk automatically detects local file paths and uploads them. urls are passed through as-is.

python
1# local file paths are automatically uploaded2result = client.run({3    "app": "heygen/voice-clone",4    "input": {5        "image": "/path/to/local/image.png",  # detected & uploaded6        "audio": "https://example.com/audio.mp3",  # url passed through7    }8})

manual upload

you can also upload files manually and use the returned url.

python
1# upload and get a hosted URL2file = client.files.upload("/path/to/file.png")3print(file.uri)  # https://cloud.inference.sh/...

4. webhooks

get notified when a task completes by providing a webhook url. when the task reaches a terminal state (completed, failed, or cancelled), a POST request is sent to your url with the task result.

python
1result = client.run({2    "app": "heygen/voice-clone",3    "input": {},4    "webhook": "https://your-server.com/webhook"5}, wait=False)

webhook payload

your endpoint receives a JSON POST with the task result:

json
1{2  "id": "task_abc123",3  "status": 9,4  "output": { ... },5  "error": "",6  "session_id": null,7  "created_at": "2024-01-15T10:30:00Z",8  "updated_at": "2024-01-15T10:30:05Z"9}
idstringtask id
statusnumberterminal status (9=completed, 10=failed, 11=cancelled)
outputobjecttask output (when completed)
errorstringerror message (when failed)
session_idstringsession id (if using sessions)
created_atstringiso timestamp
updated_atstringiso timestamp

5. schema

input

audiostring(file)*

reference recording of the voice to clone. 30-60 seconds of clean, single-speaker speech works best.

textstring

text to speak in the cloned voice. leave empty to only create the voice and return its id.

example: "Merhaba, bu benim klonlanmış sesim. Nasıl duyuluyor?"
voice_namestring

display name for the cloned voice.

default: "cloned voice"maxLength:100
languagestring

language hint for the clone (e.g. 'tr', 'en'). auto-detected if omitted.

speednumber

speech speed multiplier for the generated audio (0.5-2.0).

default: 1min:0.5max:2
remove_background_noiseboolean

clean background noise out of the reference recording before cloning.

default: true
keep_voiceboolean

keep the cloned voice in the workspace so its id can be reused later. off by default: the clone is deleted after the audio is generated, because the heygen account holds a limited number of clone slots. forced on when no text is given.

default: false

output

audiostring(file)

speech generated in the cloned voice (if text was supplied).

keptboolean*

whether the voice still exists in the workspace. false means the clone was deleted after generation and the id is no longer usable.

preview_audio_urlstring

heygen's own short preview clip of the cloned voice.

voice_idstring*

cloned voice id. usable with heygen/text-to-speech and avatar videos while the voice is kept.

voice_namestring*

display name of the cloned voice.

ready to run voice-clone?

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.