Decision 3.0 d3-lite 0.8B

The 0.8B Decision 3.0 model from vLLM Semantic Router. Send text, images and videos with typed questions (choice, score, yes/no); get a probability for every answer.

run with your agent
# install belt
$curl -fsSL https://cli.inference.sh | sh
# sign in (or set INFSH_API_KEY on a headless machine)
$belt login
# view schema & details
$belt app get vllm-sr/d3-lite
# run
$belt app run vllm-sr/d3-lite --input '{"state":"The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.","choices":[{"id":"choices_1","instructions":"instructions","options":[{"name":"options_1"},{"name":"options_2"}]}],"scores":[{"id":"scores_1","instructions":"instructions","levels":["levels_1","levels_2"]}],"nouls":[{"id":"nouls_1","instructions":"instructions"}]}'
maker
vllm-sr
category
decision
price
GPU time, billed per second
runs via
api, javascript and python sdk, belt cli, mcp
updated
2026-10-11

api reference

about

the 0.8b decision 3.0 model from vllm semantic router. send text, images and videos with typed questions (choice, score, yes/no); get a probability for every answer.

1. calling the api

install the client

the client provides a convenient way to interact with the api.

bash
1pip install inferencesh

setup your api key

set INFERENCE_API_KEY as an environment variable. get your key from settings → api keys.

bash
1export INFERENCE_API_KEY="inf_your_key"

run and get result

submit a request and wait for the final result. best for batch processing or when you don't need progress updates.

python
1import os2from inferencesh import inference34client = inference(api_key=os.environ["INFERENCE_API_KEY"])567result = client.run({8        "app": "vllm-sr/d3-lite",9        "input": {10            "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",11            "choices": [12                {13                    "id": "choices_1",14                    "instructions": "instructions",15                    "options": [16                        {17                            "name": "options_1"18                        },19                        {20                            "name": "options_2"21                        }22                    ]23                }24            ],25            "scores": [26                {27                    "id": "scores_1",28                    "instructions": "instructions",29                    "levels": [30                        "levels_1",31                        "levels_2"32                    ]33                }34            ],35            "nouls": [36                {37                    "id": "nouls_1",38                    "instructions": "instructions"39                }40            ]41        }42    })4344print(result["output"])

stream live updates

get real-time progress updates as the task runs. ideal for showing progress bars, partial results, or long-running tasks.

python
1import os2from inferencesh import inference34client = inference(api_key=os.environ["INFERENCE_API_KEY"])567# stream=True yields updates as they arrive8for update in client.run({9        "app": "vllm-sr/d3-lite",10        "input": {11            "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",12            "choices": [13                {14                    "id": "choices_1",15                    "instructions": "instructions",16                    "options": [17                        {18                            "name": "options_1"19                        },20                        {21                            "name": "options_2"22                        }23                    ]24                }25            ],26            "scores": [27                {28                    "id": "scores_1",29                    "instructions": "instructions",30                    "levels": [31                        "levels_1",32                        "levels_2"33                    ]34                }35            ],36            "nouls": [37                {38                    "id": "nouls_1",39                    "instructions": "instructions"40                }41            ]42        }43    }, stream=True):44    if update.get("progress"):45        print(f"progress: {update['progress']}%")46    if update.get("output"):47        print(f"output: {update['output']}")

2. authentication

the api uses api keys for authentication. see the authentication docs for detailed setup instructions.

3. files

file inputs are automatically handled by the sdk. you can pass local paths, urls, or base64 data.

automatic upload

the python sdk automatically detects local file paths and uploads them. urls are passed through as-is.

python
1# local file paths are automatically uploaded2result = client.run({3    "app": "vllm-sr/d3-lite",4    "input": {5        "image": "/path/to/local/image.png",  # detected & uploaded6        "audio": "https://example.com/audio.mp3",  # url passed through7    }8})

manual upload

you can also upload files manually and use the returned url.

python
1# upload and get a hosted URL2file = client.files.upload("/path/to/file.png")3print(file.uri)  # https://cloud.inference.sh/...

4. webhooks

get notified when a task completes by providing a webhook url. when the task reaches a terminal state (completed, failed, or cancelled), a POST request is sent to your url with the task result.

python
1result = client.run({2    "app": "vllm-sr/d3-lite",3    "input": {4        "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",5        "choices": [6            {7                "id": "choices_1",8                "instructions": "instructions",9                "options": [10                    {11                        "name": "options_1"12                    },13                    {14                        "name": "options_2"15                    }16                ]17            }18        ],19        "scores": [20            {21                "id": "scores_1",22                "instructions": "instructions",23                "levels": [24                    "levels_1",25                    "levels_2"26                ]27            }28        ],29        "nouls": [30            {31                "id": "nouls_1",32                "instructions": "instructions"33            }34        ]35    },36    "webhook": "https://your-server.com/webhook"37}, wait=False)

webhook payload

your endpoint receives a JSON POST with the task result:

json
1{2  "event": "task.completed",3  "timestamp": "2024-01-15T10:30:05Z",4  "data": {5    "id": "task_abc123",6    "short_id": "abc123",7    "status": 10,8    "status_text": "completed",9    "output": { ... },10    "created_at": "2024-01-15T10:30:00Z",11    "updated_at": "2024-01-15T10:30:05Z"12  }13}
eventstringtask.completed, task.failed or task.cancelled
timestampstringiso timestamp of the delivery
data.idstringtask id
data.short_idstringshort task id
data.statusnumberterminal status (10=completed, 11=failed, 12=cancelled)
data.status_textstringthe status as a word: completed, failed or cancelled
data.outputobjecttask output (when completed)
data.errorstringerror message (when failed, omitted otherwise)
data.session_idstringsession id (when the task ran in a session, omitted otherwise)
data.created_atstringiso timestamp
data.updated_atstringiso timestamp

5. schema

input

stateany

the text to evaluate: a string, or a json object / array of related context (messages, records, a policy). every question sees the same state, images and videos. may be empty when `images` or `videos` carry the content. input over the model's token limit is rejected, never truncated.

default: ""example: "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today."
imagesarray

up to 8 images (png, jpeg or webp) the questions are about, placed before the state. each is read at up to 1.6 megapixels; larger images cost more tokens and time.

videosarray

up to 4 videos (mp4, webm, mov or mkv, at most 32 mb and 5 minutes each) the questions are about, placed after the images. read at 2 frames per second, at most 32 frames per video at up to 0.2 megapixels; all videos together take at most 16,384 tokens.

choicesarray

choice questions: pick one option from a set.

scoresarray

score questions: place the state on ordered levels.

noulsarray

noul questions: probability that the answer is yes.

output

choicesobject

choice answers by question id.

input_tokensinteger

input tokens, summed over the questions. image and video tokens are included.

modelstring*

the model that answered, e.g. `d3-mini`.

noulsobject

noul answers by question id.

scoresobject

score answers by question id.

ready to run Decision 3.0 d3-lite 0.8B?

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.