apps/xai/grok-stt

Grok Speech to Text

Transcribe speech with xAI's Grok Speech to Text. Send a recording and get the transcript with word timings and speakers, or stream a microphone over a socket and read the transcript as it is spoken.

run with your agent
# install belt
$curl -fsSL https://cli.inference.sh | sh
# view schema & details
$belt app get xai/grok-stt
# run
$belt app run xai/grok-stt

api reference

about

transcribe speech with xai's grok speech to text. send a recording and get the transcript with word timings and speakers, or stream a microphone over a socket and read the transcript as it is spoken.

1. calling the api

install the client

the client provides a convenient way to interact with the api.

bash
1pip install inferencesh

setup your api key

set INFERENCE_API_KEY as an environment variable. get your key from settings → api keys.

bash
1export INFERENCE_API_KEY="inf_your_key"

run and get result

submit a request and wait for the final result. best for batch processing or when you don't need progress updates.

python
1from inferencesh import inference23client = inference()456result = client.run({7        "app": "xai/grok-stt",8        "input": {}9    })1011print(result["output"])

stream live updates

get real-time progress updates as the task runs. ideal for showing progress bars, partial results, or long-running tasks.

python
1from inferencesh import inference23client = inference()456# stream=True yields updates as they arrive7for update in client.run({8        "app": "xai/grok-stt",9        "input": {}10    }, stream=True):11    if update.get("progress"):12        print(f"progress: {update['progress']}%")13    if update.get("output"):14        print(f"output: {update['output']}")

2. authentication

the api uses api keys for authentication. see the authentication docs for detailed setup instructions.

3. files

file inputs are automatically handled by the sdk. you can pass local paths, urls, or base64 data.

automatic upload

the python sdk automatically detects local file paths and uploads them. urls are passed through as-is.

python
1# local file paths are automatically uploaded2result = client.run({3    "app": "xai/grok-stt",4    "input": {5        "image": "/path/to/local/image.png",  # detected & uploaded6        "audio": "https://example.com/audio.mp3",  # url passed through7    }8})

manual upload

you can also upload files manually and use the returned url.

python
1# upload and get a hosted URL2file = client.files.upload("/path/to/file.png")3print(file.uri)  # https://cloud.inference.sh/...

4. webhooks

get notified when a task completes by providing a webhook url. when the task reaches a terminal state (completed, failed, or cancelled), a POST request is sent to your url with the task result.

python
1result = client.run({2    "app": "xai/grok-stt",3    "input": {},4    "webhook": "https://your-server.com/webhook"5}, wait=False)

webhook payload

your endpoint receives a JSON POST with the task result:

json
1{2  "id": "task_abc123",3  "status": 9,4  "output": { ... },5  "error": "",6  "session_id": null,7  "created_at": "2024-01-15T10:30:00Z",8  "updated_at": "2024-01-15T10:30:05Z"9}
idstring— task id
statusnumber— terminal status (9=completed, 10=failed, 11=cancelled)
outputobject— task output (when completed)
errorstring— error message (when failed)
session_idstring— session id (if using sessions)
created_atstring— iso timestamp
updated_atstring— iso timestamp

5. schema

input

audiostring(file)*

the recording: wav, mp3, ogg, opus, flac, aac, m4a, mp4 or mkv, up to 500 mb

languagestring

what is spoken. speech is transcribed in any supported language without it; naming it writes numbers, currencies and units the way they are written ("$167,983.15").

options:"ar""cs""da""nl""en""fil""fr""de""hi""id""it""ja""ko""mk""ms""fa""pl""pt""ro""ru""es""sv""th""tr""vi"
diarizeboolean

tell speakers apart: each word carries a speaker number

default: false
keytermsarray

names and terms the transcript should prefer, up to 100 of at most 50 characters each

filler_wordsboolean

keep fillers such as "um" and "uh"

default: false

output

durationnumber

length of the audio, in seconds

languagestring

the language xai detected, such as en or es-mx

textstring*

the transcript

wordsarray

every word with its timing

ready to run Grok Speech to Text?

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.