
Decision 3.0 d3-edge 0.6B
The smallest and fastest Decision 3.0 model from vLLM Semantic Router. Send text, images and videos with typed questions (choice, score, yes/no); get a probability for every answer.
api reference
about
the smallest and fastest decision 3.0 model from vllm semantic router. send text, images and videos with typed questions (choice, score, yes/no); get a probability for every answer.
1. calling the api
install the client
the client provides a convenient way to interact with the api.
1pip install inferenceshsetup your api key
set INFERENCE_API_KEY as an environment variable. get your key from settings → api keys.
1export INFERENCE_API_KEY="inf_your_key"run and get result
submit a request and wait for the final result. best for batch processing or when you don't need progress updates.
1import os2from inferencesh import inference34client = inference(api_key=os.environ["INFERENCE_API_KEY"])567result = client.run({8 "app": "vllm-sr/d3-edge",9 "input": {10 "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",11 "choices": [12 {13 "id": "choices_1",14 "instructions": "instructions",15 "options": [16 {17 "name": "options_1"18 },19 {20 "name": "options_2"21 }22 ]23 }24 ],25 "scores": [26 {27 "id": "scores_1",28 "instructions": "instructions",29 "levels": [30 "levels_1",31 "levels_2"32 ]33 }34 ],35 "nouls": [36 {37 "id": "nouls_1",38 "instructions": "instructions"39 }40 ]41 }42 })4344print(result["output"])stream live updates
get real-time progress updates as the task runs. ideal for showing progress bars, partial results, or long-running tasks.
1import os2from inferencesh import inference34client = inference(api_key=os.environ["INFERENCE_API_KEY"])567# stream=True yields updates as they arrive8for update in client.run({9 "app": "vllm-sr/d3-edge",10 "input": {11 "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",12 "choices": [13 {14 "id": "choices_1",15 "instructions": "instructions",16 "options": [17 {18 "name": "options_1"19 },20 {21 "name": "options_2"22 }23 ]24 }25 ],26 "scores": [27 {28 "id": "scores_1",29 "instructions": "instructions",30 "levels": [31 "levels_1",32 "levels_2"33 ]34 }35 ],36 "nouls": [37 {38 "id": "nouls_1",39 "instructions": "instructions"40 }41 ]42 }43 }, stream=True):44 if update.get("progress"):45 print(f"progress: {update['progress']}%")46 if update.get("output"):47 print(f"output: {update['output']}")2. authentication
the api uses api keys for authentication. see the authentication docs for detailed setup instructions.
3. files
file inputs are automatically handled by the sdk. you can pass local paths, urls, or base64 data.
automatic upload
the python sdk automatically detects local file paths and uploads them. urls are passed through as-is.
1# local file paths are automatically uploaded2result = client.run({3 "app": "vllm-sr/d3-edge",4 "input": {5 "image": "/path/to/local/image.png", # detected & uploaded6 "audio": "https://example.com/audio.mp3", # url passed through7 }8})4. webhooks
get notified when a task completes by providing a webhook url. when the task reaches a terminal state (completed, failed, or cancelled), a POST request is sent to your url with the task result.
1result = client.run({2 "app": "vllm-sr/d3-edge",3 "input": {4 "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",5 "choices": [6 {7 "id": "choices_1",8 "instructions": "instructions",9 "options": [10 {11 "name": "options_1"12 },13 {14 "name": "options_2"15 }16 ]17 }18 ],19 "scores": [20 {21 "id": "scores_1",22 "instructions": "instructions",23 "levels": [24 "levels_1",25 "levels_2"26 ]27 }28 ],29 "nouls": [30 {31 "id": "nouls_1",32 "instructions": "instructions"33 }34 ]35 },36 "webhook": "https://your-server.com/webhook"37}, wait=False)webhook payload
your endpoint receives a JSON POST with the task result:
1{2 "event": "task.completed",3 "timestamp": "2024-01-15T10:30:05Z",4 "data": {5 "id": "task_abc123",6 "short_id": "abc123",7 "status": 10,8 "status_text": "completed",9 "output": { ... },10 "created_at": "2024-01-15T10:30:00Z",11 "updated_at": "2024-01-15T10:30:05Z"12 }13}5. schema
input
the text to evaluate: a string, or a json object / array of related context (messages, records, a policy). every question sees the same state, images and videos. may be empty when `images` or `videos` carry the content. input over the model's token limit is rejected, never truncated.
up to 8 images (png, jpeg or webp) the questions are about, placed before the state. each is read at up to 1.6 megapixels; larger images cost more tokens and time.
up to 4 videos (mp4, webm, mov or mkv, at most 32 mb and 5 minutes each) the questions are about, placed after the images. read at 2 frames per second, at most 32 frames per video at up to 0.2 megapixels; all videos together take at most 16,384 tokens.
choice questions: pick one option from a set.
score questions: place the state on ordered levels.
noul questions: probability that the answer is yes.
Decision 3.0 d3 27B
vllm-sr/d3
GPU time, billed per second
Decision 3.0 d3-flash 9B
vllm-sr/d3-flash
GPU time, billed per second
Decision 3.0 d3-mini 4B
vllm-sr/d3-mini
GPU time, billed per second
Decision 3.0 d3-nano 2B
vllm-sr/d3-nano
GPU time, billed per second
Decision 3.0 d3-lite 0.8B
vllm-sr/d3-lite
GPU time, billed per second
ready to run Decision 3.0 d3-edge 0.6B?
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.