# Gemini 3.1 Pro Preview

> Run Gemini 3.1 Pro Preview by openrouter online or via API. From $2.00 per 1M tokens, no setup.

**URL:** https://inference.sh/apps/openrouter/gemini-3-1-pro-preview
**Last updated:** 2026-10-11

---

Google's frontier reasoning model, with stronger coding and agentic reliability. 1M context; text, image, file, audio and video input.

- app: `openrouter/gemini-3-1-pro-preview`
- maker: [openrouter](https://inference.sh/apps/openrouter.md)
- category: [chat](https://inference.sh/apps/category/chat.md)
- tasks: [reasoning](https://inference.sh/apps/tag/reasoning.md)
- price: $2.00/M input, $12.00/M output; >200K ctx: $4.00/M in, $18.00/M out
- runs via: api, javascript and python sdk, belt cli, mcp

## gemini-3-1-pro-preview API Guide

### About
google's frontier reasoning model, with stronger coding and agentic reliability. 1m context; text, image, file, audio and video input.

### 1. Calling the API

#### Install the Client

```bash
pip install inferencesh
```

#### Setup Your API Key
Set `INFERENCE_API_KEY` as an environment variable. Get your key from settings → api keys.

```bash
export INFERENCE_API_KEY="inf_your_key"
```

#### Run and Get Result

```python
from inferencesh import inference

client = inference()


result = client.run({
        "app": "openrouter/gemini-3-1-pro-preview@4kqtv2x5",
        "input": {}
    })

print(result["output"])
```

#### Stream Live Updates

```python
from inferencesh import inference

client = inference()


## stream=True yields updates as they arrive
for update in client.run({
        "app": "openrouter/gemini-3-1-pro-preview@4kqtv2x5",
        "input": {}
    }, stream=True):
    if update.get("progress"):
        print(f"progress: {update['progress']}%")
    if update.get("output"):
        print(f"output: {update['output']}")
```

### 2. Authentication
The API uses API keys for authentication. See the authentication docs for detailed setup instructions.

### 3. Files

#### Automatic Upload

```python
## local file paths are automatically uploaded
result = client.run({
    "app": "openrouter/gemini-3-1-pro-preview@4kqtv2x5",
    "input": {
        "image": "/path/to/local/image.png",  # detected & uploaded
        "audio": "https://example.com/audio.mp3",  # url passed through
    }
})
```

#### Manual Upload

```python
## upload and get a hosted URL
file = client.files.upload("/path/to/file.png")
print(file.uri)  # https://cloud.inference.sh/...
```

### 4. Schema

#### Input
- **reasoning_effort** (string): Reasoning effort. This model always reasons; it cannot be turned off.
- **reasoning_exclude** (boolean): Exclude reasoning tokens from response
- **context_size** (integer): The context size for the model.
- **temperature** (number): Temperature
- **top_p** (number): Top P
- **top_k** (integer): Top K
- **min_p** (number): Min P
- **frequency_penalty** (number): Frequency Penalty
- **presence_penalty** (number): Presence Penalty
- **repetition_penalty** (number): Repetition Penalty
- **seed** (integer): Seed
- **stop** (array): Stop
- **max_tokens** (integer): Max Tokens
- **reasoning_max_tokens** (integer): Reasoning Max Tokens
- **system_prompt** (string): System Prompt
- **tools** (array): Tools
- **tool_choice** (object): ToolChoice
- **response_format** (object): ResponseFormat
- **context** (array): Context
- **role** (string): ChatMessageRole
- **text** (string): Text
- **reasoning** (string): Reasoning
- **attachments** (array): Attachments
- **images** (array): Images
- **files** (array): Files
- **tool_call_id** (string): Tool Call Id

#### Output
- **images** (array): Images
- **reasoning** (string): Reasoning
- **response** (string): Response
- **tool_calls** (array): Tool Calls
- **usage** (object): LLMUsage

## other chat apps

- [gemini-3-7-flash](https://inference.sh/apps/google/gemini-3-7-flash.md): $1.50/M input, $7.50/M output
- [gemini-3-8-flash](https://inference.sh/apps/google/gemini-3-8-flash.md): $1.50/M input, $7.50/M output
- [DeepSeek V4.1 Flash](https://inference.sh/apps/melious/deepseek-v4-1-flash.md): $0.2273/M input, $0.0114/M cached input, $1.1367/M output
- [MiniMax M3](https://inference.sh/apps/melious/minimax-m3.md): $0.4547/M input, $0.1137/M cached input, $2.2734/M output
- [GLM 5.3](https://inference.sh/apps/melious/glm-5-3.md): $1.1367/M input, $0.2273/M cached input, $3.4101/M output

## faq

### how much does Gemini 3.1 Pro Preview cost?

$2.00/M input, $12.00/M output; >200K ctx: $4.00/M in, $18.00/M out. you pay per run from your inference shell credits. no subscription is required.

### how do I run Gemini 3.1 Pro Preview via api?

call openrouter/gemini-3-1-pro-preview with the javascript or python sdk, the rest api, or `belt app run openrouter/gemini-3-1-pro-preview` from a terminal. the api reference on this page lists every input.

### who makes Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is made by openrouter and runs on inference shell. every openrouter app is listed at https://inference.sh/apps/openrouter.

### can my agent use Gemini 3.1 Pro Preview?

yes. agents call it as a tool through the inference shell mcp server (https://api.inference.sh/mcp), the belt cli or the sdk, and pay per run like you do.

### what are the alternatives to Gemini 3.1 Pro Preview?

other chat apps on inference shell include gemini-3-7-flash, gemini-3-8-flash, DeepSeek V4.1 Flash, MiniMax M3. they run through the same api, so switching is a one-line change.
