# Decision 3.0 d3-nano 2B

> Run Decision 3.0 d3-nano 2B by vllm-sr online or via API. Pay per run, no setup.

**URL:** https://inference.sh/apps/vllm-sr/d3-nano
**Last updated:** 2026-10-11

---

The 2B Decision 3.0 model from vLLM Semantic Router. Send text, images and videos with typed questions (choice, score, yes/no); get a probability for every answer.

- app: `vllm-sr/d3-nano`
- maker: [vllm-sr](https://inference.sh/apps/vllm-sr.md)
- category: [decision](https://inference.sh/apps/category/decision.md)
- tasks: [classification](https://inference.sh/apps/tag/classification.md), [routing](https://inference.sh/apps/tag/routing.md), [moderation](https://inference.sh/apps/tag/moderation.md), [vision](https://inference.sh/apps/tag/vision.md), [video](https://inference.sh/apps/tag/video.md), [open weights](https://inference.sh/apps/tag/open-weights.md)
- price: GPU time, billed per second
- runs via: api, javascript and python sdk, belt cli, mcp

## d3-nano API Guide

### About
the 2b decision 3.0 model from vllm semantic router. send text, images and videos with typed questions (choice, score, yes/no); get a probability for every answer.

### 1. Calling the API

#### Install the Client

```bash
pip install inferencesh
```

#### Setup Your API Key
Set `INFERENCE_API_KEY` as an environment variable. Get your key from settings → api keys.

```bash
export INFERENCE_API_KEY="inf_your_key"
```

#### Run and Get Result

```python
import os
from inferencesh import inference

client = inference(api_key=os.environ["INFERENCE_API_KEY"])


result = client.run({
        "app": "vllm-sr/d3-nano@0q17m8t2",
        "input": {
            "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
            "choices": [
                {
                    "id": "choices_1",
                    "instructions": "instructions",
                    "options": [
                        {
                            "name": "options_1"
                        },
                        {
                            "name": "options_2"
                        }
                    ]
                }
            ],
            "scores": [
                {
                    "id": "scores_1",
                    "instructions": "instructions",
                    "levels": [
                        "levels_1",
                        "levels_2"
                    ]
                }
            ],
            "nouls": [
                {
                    "id": "nouls_1",
                    "instructions": "instructions"
                }
            ]
        }
    })

print(result["output"])
```

#### Stream Live Updates

```python
import os
from inferencesh import inference

client = inference(api_key=os.environ["INFERENCE_API_KEY"])


## stream=True yields updates as they arrive
for update in client.run({
        "app": "vllm-sr/d3-nano@0q17m8t2",
        "input": {
            "state": "The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
            "choices": [
                {
                    "id": "choices_1",
                    "instructions": "instructions",
                    "options": [
                        {
                            "name": "options_1"
                        },
                        {
                            "name": "options_2"
                        }
                    ]
                }
            ],
            "scores": [
                {
                    "id": "scores_1",
                    "instructions": "instructions",
                    "levels": [
                        "levels_1",
                        "levels_2"
                    ]
                }
            ],
            "nouls": [
                {
                    "id": "nouls_1",
                    "instructions": "instructions"
                }
            ]
        }
    }, stream=True):
    if update.get("progress"):
        print(f"progress: {update['progress']}%")
    if update.get("output"):
        print(f"output: {update['output']}")
```

### 2. Authentication
The API uses API keys for authentication. See the authentication docs for detailed setup instructions.

### 3. Files

#### Automatic Upload

```python
## local file paths are automatically uploaded
result = client.run({
    "app": "vllm-sr/d3-nano@0q17m8t2",
    "input": {
        "image": "/path/to/local/image.png",  # detected & uploaded
        "audio": "https://example.com/audio.mp3",  # url passed through
    }
})
```

#### Manual Upload

```python
## upload and get a hosted URL
file = client.files.upload("/path/to/file.png")
print(file.uri)  # https://cloud.inference.sh/...
```

### 4. Schema

#### Input
- **state** (any): The text to evaluate: a string, or a JSON object / array of related context (messages, records, a policy). Every question sees the same state, images and videos. May be empty when `images` or `videos` carry the content. Input over the model's token limit is rejected, never truncated.
- **images** (array): Up to 8 images (PNG, JPEG or WebP) the questions are about, placed before the state. Each is read at up to 1.6 megapixels; larger images cost more tokens and time.
- **videos** (array): Up to 4 videos (MP4, WebM, MOV or MKV, at most 32 MB and 5 minutes each) the questions are about, placed after the images. Read at 2 frames per second, at most 32 frames per video at up to 0.2 megapixels; all videos together take at most 16,384 tokens.
- **choices** (array): Choice questions: pick one option from a set.
- **scores** (array): Score questions: place the state on ordered levels.
- **nouls** (array): Noul questions: probability that the answer is yes.

#### Output
- **choices** (object): Choice answers by question id.
- **input_tokens** (integer): Input tokens, summed over the questions. Image and video tokens are included.
- **model** (string): The model that answered, e.g. `d3-mini`.
- **nouls** (object): Noul answers by question id.
- **scores** (object): Score answers by question id.

## other decision apps

- [Decision 3.0 d3 27B](https://inference.sh/apps/vllm-sr/d3.md): GPU time, billed per second
- [Decision 3.0 d3-flash 9B](https://inference.sh/apps/vllm-sr/d3-flash.md): GPU time, billed per second
- [Decision 3.0 d3-mini 4B](https://inference.sh/apps/vllm-sr/d3-mini.md): GPU time, billed per second
- [Decision 3.0 d3-lite 0.8B](https://inference.sh/apps/vllm-sr/d3-lite.md): GPU time, billed per second
- [Decision 3.0 d3-edge 0.6B](https://inference.sh/apps/vllm-sr/d3-edge.md): GPU time, billed per second

## faq

### how much does Decision 3.0 d3-nano 2B cost?

GPU time, billed per second. you pay per run from your inference shell credits. no subscription is required.

### how do I run Decision 3.0 d3-nano 2B via api?

call vllm-sr/d3-nano with the javascript or python sdk, the rest api, or `belt app run vllm-sr/d3-nano` from a terminal. the api reference on this page lists every input.

### who makes Decision 3.0 d3-nano 2B?

Decision 3.0 d3-nano 2B is made by vllm-sr and runs on inference shell. every vllm-sr app is listed at https://inference.sh/apps/vllm-sr.

### can my agent use Decision 3.0 d3-nano 2B?

yes. agents call it as a tool through the inference shell mcp server (https://api.inference.sh/mcp), the belt cli or the sdk, and pay per run like you do.

### what are the alternatives to Decision 3.0 d3-nano 2B?

other decision apps on inference shell include Decision 3.0 d3 27B, Decision 3.0 d3-flash 9B, Decision 3.0 d3-mini 4B, Decision 3.0 d3-lite 0.8B. they run through the same api, so switching is a one-line change.
