# Decision 2.0 Sol 2B Reasoning

> Run Decision 2.0 Sol 2B Reasoning by vllm-sr online or via API. Pay per run, no setup.

**URL:** https://inference.sh/apps/vllm-sr/decision-2-0-sol-2b-reasoning
**Last updated:** 2026-10-08

---

The reasoning variant of Decision 2.0 Sol 2B, stronger on multi-step problems (arithmetic, code execution, causal and logical questions) at the same speed from vLLM Semantic Router. Send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 1.88B parameters, 16,384-token input, about 7.5 ms per question.

- app: `vllm-sr/decision-2-0-sol-2b-reasoning`
- maker: [vllm-sr](https://inference.sh/apps/vllm-sr.md)
- category: [decision](https://inference.sh/apps/category/decision.md)
- tasks: [classification](https://inference.sh/apps/tag/classification.md), [routing](https://inference.sh/apps/tag/routing.md), [moderation](https://inference.sh/apps/tag/moderation.md), [open weights](https://inference.sh/apps/tag/open-weights.md)
- price: GPU time, billed per second
- runs via: api, javascript and python sdk, belt cli, mcp

## decision-2-0-sol-2b-reasoning API Guide

### About
the reasoning variant of decision 2.0 sol 2b, stronger on multi-step problems (arithmetic, code execution, causal and logical questions) at the same speed from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 1.88b parameters, 16,384-token input, about 7.5 ms per question.

### 1. Calling the API

#### Install the Client

```bash
pip install inferencesh
```

#### Setup Your API Key
Set `INFERENCE_API_KEY` as an environment variable. Get your key from settings → api keys.

```bash
export INFERENCE_API_KEY="inf_your_key"
```

#### Run and Get Result

```python
from inferencesh import inference

client = inference()


result = client.run({
        "app": "vllm-sr/decision-2-0-sol-2b-reasoning@3s6qh0mt",
        "input": {}
    })

print(result["output"])
```

#### Stream Live Updates

```python
from inferencesh import inference

client = inference()


## stream=True yields updates as they arrive
for update in client.run({
        "app": "vllm-sr/decision-2-0-sol-2b-reasoning@3s6qh0mt",
        "input": {}
    }, stream=True):
    if update.get("progress"):
        print(f"progress: {update['progress']}%")
    if update.get("output"):
        print(f"output: {update['output']}")
```

### 2. Authentication
The API uses API keys for authentication. See the authentication docs for detailed setup instructions.

### 3. Files

#### Automatic Upload

```python
## local file paths are automatically uploaded
result = client.run({
    "app": "vllm-sr/decision-2-0-sol-2b-reasoning@3s6qh0mt",
    "input": {
        "image": "/path/to/local/image.png",  # detected & uploaded
        "audio": "https://example.com/audio.mp3",  # url passed through
    }
})
```

#### Manual Upload

```python
## upload and get a hosted URL
file = client.files.upload("/path/to/file.png")
print(file.uri)  # https://cloud.inference.sh/...
```

### 4. Schema

#### Input
- **state** (any) *required*: The content to evaluate: a string, or a JSON object / array of related context (messages, records, a policy). Text only. Every question sees the same state. Input over the model's token limit is rejected, never truncated.
- **choices** (array): Choice questions: pick one option from a set.
- **scores** (array): Score questions: place the state on ordered levels.
- **nouls** (array): Noul questions: probability that the answer is yes.

#### Output
- **choices** (object): Choice answers by question id.
- **input_tokens** (integer): Input tokens, summed over the questions.
- **model** (string): The model that answered, e.g. `Decision-2.0-Sol-2B`.
- **nouls** (object): Noul answers by question id.
- **scores** (object): Score answers by question id.

## other decision apps

- [OpenAI Decisions](https://inference.sh/apps/openai/decisions.md): $0.10/M input, output free
- [d1-omni-600M](https://inference.sh/apps/liquid/d1-omni-600m.md): from $0.001/sec
- [d1-3B](https://inference.sh/apps/liquid/d1-3b.md): from $0.001/sec
- [Perplexity Decider v1.1 27B](https://inference.sh/apps/perplexity/decider-v1-1-27b.md): from $0.001/sec
- [Decision 2.0 Vega 27B](https://inference.sh/apps/vllm-sr/decision-2-0-vega-27b.md): from $0.001/sec

## faq

### how much does Decision 2.0 Sol 2B Reasoning cost?

GPU time, billed per second. you pay per run from your inference shell credits. no subscription is required.

### how do I run Decision 2.0 Sol 2B Reasoning via api?

call vllm-sr/decision-2-0-sol-2b-reasoning with the javascript or python sdk, the rest api, or `belt app run vllm-sr/decision-2-0-sol-2b-reasoning` from a terminal. the api reference on this page lists every input.

### who makes Decision 2.0 Sol 2B Reasoning?

Decision 2.0 Sol 2B Reasoning is made by vllm-sr and runs on inference shell. every vllm-sr app is listed at https://inference.sh/apps/vllm-sr.

### can my agent use Decision 2.0 Sol 2B Reasoning?

yes. agents call it as a tool through the inference shell mcp server (https://api.inference.sh/mcp), the belt cli or the sdk, and pay per run like you do.

### what are the alternatives to Decision 2.0 Sol 2B Reasoning?

other decision apps on inference shell include OpenAI Decisions, d1-omni-600M, d1-3B, Perplexity Decider v1.1 27B. they run through the same api, so switching is a one-line change.
