
grok-voice
Talk with Grok. Stream microphone audio over a socket and hear xAI's speech-to-speech model answer as it speaks, with both sides transcribed; the conversation is the task's output.
api reference
about
talk with grok. stream microphone audio over a socket and hear xai's speech-to-speech model answer as it speaks, with both sides transcribed; the conversation is the task's output.
1. calling the api
install the client
the client provides a convenient way to interact with the api.
1pip install inferenceshsetup your api key
set INFERENCE_API_KEY as an environment variable. get your key from settings → api keys.
1export INFERENCE_API_KEY="inf_your_key"run and get result
submit a request and wait for the final result. best for batch processing or when you don't need progress updates.
1from inferencesh import inference23client = inference()456result = client.run({7 "app": "xai/grok-voice",8 "input": {}9 })1011print(result["output"])stream live updates
get real-time progress updates as the task runs. ideal for showing progress bars, partial results, or long-running tasks.
1from inferencesh import inference23client = inference()456# stream=True yields updates as they arrive7for update in client.run({8 "app": "xai/grok-voice",9 "input": {}10 }, stream=True):11 if update.get("progress"):12 print(f"progress: {update['progress']}%")13 if update.get("output"):14 print(f"output: {update['output']}")2. authentication
the api uses api keys for authentication. see the authentication docs for detailed setup instructions.
3. files
file inputs are automatically handled by the sdk. you can pass local paths, urls, or base64 data.
automatic upload
the python sdk automatically detects local file paths and uploads them. urls are passed through as-is.
1# local file paths are automatically uploaded2result = client.run({3 "app": "xai/grok-voice",4 "input": {5 "image": "/path/to/local/image.png", # detected & uploaded6 "audio": "https://example.com/audio.mp3", # url passed through7 }8})4. webhooks
get notified when a task completes by providing a webhook url. when the task reaches a terminal state (completed, failed, or cancelled), a POST request is sent to your url with the task result.
1result = client.run({2 "app": "xai/grok-voice",3 "input": {},4 "webhook": "https://your-server.com/webhook"5}, wait=False)webhook payload
your endpoint receives a JSON POST with the task result:
1{2 "id": "task_abc123",3 "status": 9,4 "output": { ... },5 "error": "",6 "session_id": null,7 "created_at": "2024-01-15T10:30:00Z",8 "updated_at": "2024-01-15T10:30:05Z"9}5. schema
input
microphone audio, a frame every 20 ms or so
typed messages: text from the user, or words for the assistant
system prompt. grok's voice models take plain instructions; workarounds written for other models are unnecessary.
a built-in voice; `voices` lists them with what xai says about each
the id of a voice cloned with xai's custom voices api. set, it is used instead of voice.
grok-voice-latest follows the newest model; pin a versioned name such as grok-voice-think-fast-2.0 for stability
whether the model thinks before it answers. 'none' answers faster.
bcp-47 hint for what the user speaks, such as en, ja or es-mx (spanish and portuguese need a region). left empty it is detected.
playback speed of the assistant's voice
silence that ends the user's turn, in ms. left empty grok decides.
let the assistant search the web
let the assistant search x
output
what the assistant is saying, as it says it
the assistant's voice
the conversation so far, a message per turn
true while the conversation goes on; false on the result
how long the session has run
what grok hears the user saying, refined as they speak
ready to run grok-voice?
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.