apps/reasoning

reasoning

58 reasoning models and apps on inference shell.run via API, SDK, or belt CLI.

58apps
1category

chat

58
gemini-3-7-flash

gemini-3-7-flash

google/gemini-3-7-flash

gemini 3.7 flash via vertex ai — improved software engineering, web development and agentic workflows over 3.6 flash. 1m token context window.

chat
gemini-3-8-flash

gemini-3-8-flash

google/gemini-3-8-flash

gemini 3.8 flash via vertex ai — google's most capable flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows. 1m token context window.

chat
DeepSeek V4.1 Flash

DeepSeek V4.1 Flash

melious/deepseek-v4-1-flash

deepseek v4.1 flash — 552b multimodal moe with 1m context, image input and controllable reasoning. frontier agentic and coding performance. via the melious api.

chat
MiniMax M3

MiniMax M3

melious/minimax-m3

minimax m3 — 428b multimodal moe with 1m context and image input for long-horizon agentic, coding and cowork tasks. via the melious api.

chat
GLM 5.3

GLM 5.3

melious/glm-5-3

glm 5.3 — zai's 744b moe post-trained for complex coding and long-horizon agentic tasks. 1m context, thinking on by default. via the melious api.

chat
Kimi K3

Kimi K3

melious/kimi-k3

kimi k3 — moonshot's 2.8t open-weight multimodal agentic model with native vision and 1m context for long-horizon coding and reasoning. via the melious api.

chat
Qwen 3.6 27B

Qwen 3.6 27B

melious/qwen3-6-27b

qwen 3.6 27b — dense 27b vision-language model with 262k context and thinking mode. strong coding and multimodal reasoning, apache 2.0. via the melious api.

chat
GLM 5.3 Flash

GLM 5.3 Flash

melious/glm-5-3-flash

glm 5.3 flash — zai's 320b moe (18b active), the first natively multimodal glm-5 model. 1m context, image input, low cost. via the melious api.

chat
DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813

melious/deepseek-v4-pro-0813

deepseek v4 pro 0813 — 1.6t moe with 1m context and low/high/max reasoning effort for agentic and coding work. text-only. via the melious api.

chat
GLM 5.2

GLM 5.2

melious/glm-5-2

glm 5.2 — zai's 744b moe with 1m context and hybrid reasoning for coding and agentic tasks. mit license. via the melious api.

chat
DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731

melious/deepseek-v4-flash-0731

deepseek v4 flash 0731 — 304b moe with 1m context and low/high/max reasoning effort. strong agentic coding and tool use. text-only. via the melious api.

chat
grok-4-7

grok-4-7

xai/grok-4-7

grok 4.7 — xai's most capable model for chat, code and agents. 500k context, reasoning effort low to xhigh, image input and tool use via the xai responses api. direct api.

chat
grok-build-0-1

grok-build-0-1

xai/grok-build-0-1

grok build 0.1 — xai's agentic coding model. 256k context, reasoning, image input and tool use via the xai responses api. direct api.

chat
grok-4-3

grok-4-3

xai/grok-4-3

grok 4.3 — xai's fast, low-cost model. 1m context, reasoning effort from none to xhigh, image input and tool use via the xai responses api. direct api.

chat
GPT-5.6 Sol

GPT-5.6 Sol

openai/gpt-5-6-sol

gpt-5.6 sol — openai's flagship gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

chat
GPT-5.6 Luna

GPT-5.6 Luna

openai/gpt-5-6-luna

gpt-5.6 luna — openai's fast, low-cost gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

chat
GPT-5.6 Terra

GPT-5.6 Terra

openai/gpt-5-6-terra

gpt-5.6 terra — openai's balanced gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

chat
GPT-6 Astra

GPT-6 Astra

openai/gpt-6-astra

gpt-6 astra — openai's frontier model. 1m context, 128k output, adaptive reasoning (low to max), vision, file input and tool use via the responses api. direct api.

chat
Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite

google/gemini-2-5-flash-lite

gemini 2.5 flash-lite via vertex ai — most cost-efficient gemini 2.5 model optimized for high-throughput, low-latency workloads. 1m token context window.

chat
Gemini 2.5 Flash

Gemini 2.5 Flash

google/gemini-2-5-flash

gemini 2.5 flash via vertex ai — fast, cost-efficient thinking model with strong reasoning, coding, and multimodal performance. 1m token context window.

chat
Gemini 2.5 Pro

Gemini 2.5 Pro

google/gemini-2-5-pro

gemini 2.5 pro via vertex ai — google's most capable thinking model with enhanced reasoning, coding, math, and science performance. 1m token context window.

chat
MiniMax M2.7

MiniMax M2.7

minimax/m-2-7

minimax-m2.7 — large language model with enhanced reasoning, image understanding, and file processing. 200k context. direct minimax api.

chat
MiniMax M3

MiniMax M3

minimax/m3

minimax-m3 — frontier multimodal model with 1m context window. text, image, and video inputs. advanced coding, reasoning, and long-horizon agentic tasks. direct minimax api.

chat
MiniMax M2.7 Highspeed

MiniMax M2.7 Highspeed

minimax/m-2-7-highspeed

minimax-m2.7-highspeed — m2.7 performance with significantly accelerated inference. 200k context. direct minimax api.

chat
Claude Opus 5

Claude Opus 5

anthropic/claude-opus-5

claude opus 5 — anthropic's frontier opus for complex agentic coding and long-horizon work. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

chat
Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite

google/gemini-3-5-flash-lite

gemini 3.5 flash-lite via vertex ai — fastest, most cost-effective 3.5-class model at 350 output tokens/s. built for high-throughput agentic workflows.

chat
Gemini 3.6 Flash

Gemini 3.6 Flash

google/gemini-3-6-flash

gemini 3.6 flash via vertex ai — efficient workhorse model with improved coding, knowledge work, multimodal performance, and 17% fewer output tokens than 3.5 flash.

chat
Claude Sonnet 5

Claude Sonnet 5

anthropic/claude-sonnet-5

claude sonnet 5 — frontier sonnet with near-opus performance. 1m context, vision, extended thinking, tool use. direct api.

chat
Claude Opus 4.8

Claude Opus 4.8

anthropic/claude-opus-4-8

claude opus 4.8 — anthropic's most capable opus model. 1m context, 128k output, vision, extended thinking, tool use. direct api.

chat
Claude Mythos 5

Claude Mythos 5

anthropic/claude-mythos-5

claude mythos 5 — project glasswing. successor to claude mythos preview. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

chat
Claude Fable 5

Claude Fable 5

anthropic/claude-fable-5

claude fable 5 — anthropic's most capable widely released model. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

chat
Qwen3 32B

Qwen3 32B

openrouter/qwen3-32b

qwen3 32b - powerful dense language model with reasoning and tool use capabilities via openrouter

chat
Qwen3 8B

Qwen3 8B

openrouter/qwen3-8b

qwen3 8b - efficient dense language model with reasoning and tool use capabilities via openrouter

chat
Claude Haiku 4.5

Claude Haiku 4.5

anthropic/claude-haiku-4-5

claude haiku 4.5 — fastest and most affordable claude. 200k context, 64k output, vision, extended thinking, tool use. direct api.

chat
Claude Sonnet 4.5

Claude Sonnet 4.5

anthropic/claude-sonnet-4-5

claude sonnet 4.5 — previous generation sonnet. 200k context, 64k output, vision, extended thinking, tool use. direct api.

chat
Claude Sonnet 4.6

Claude Sonnet 4.6

anthropic/claude-sonnet-4-6

claude sonnet 4.6 — best balance of speed and intelligence. 1m context, 64k output, vision, extended thinking, tool use. direct api.

chat
Claude Opus 4.6

Claude Opus 4.6

anthropic/claude-opus-4-6

claude opus 4.6 — previous generation opus. 1m context, 128k output, vision, extended thinking, tool use. direct api.

chat
Claude Opus 4.7

Claude Opus 4.7

anthropic/claude-opus-4-7

claude opus 4.7 — anthropic's most capable model. 1m context, 128k output, vision, extended thinking, tool use. direct api.

chat
Claude Sonnet 4.6

Claude Sonnet 4.6

openrouter/claude-sonnet-46

sonnet 4.

chat
Hunyuan Turbos Hy3 Preview

Hunyuan Turbos Hy3 Preview

openrouter/hy3-preview

hy3 preview is a high-efficiency mixture-of-experts model from tencent designed for agentic workflows and production use.

chat
Gemini 3 Flash Preview

Gemini 3 Flash Preview

openrouter/gemini-3-flash-preview

gemini 3 flash preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

chat
Kimi K2.6

Kimi K2.6

openrouter/kimi-k26

kimi k2.

chat
Claude Opus 4.7

Claude Opus 4.7

openrouter/claude-opus-47

opus 4.

chat
Kimi K2.5

Kimi K2.5

openrouter/kimi-k25

kimi k2.5 is the latest model from moonshot ai, building on k2 with enhanced reasoning and capabilities.

chat
GPT-OSS Safeguard 20B

GPT-OSS Safeguard 20B

openrouter/gpt-oss-safeguard-20b

gpt-oss safeguard 20b

chat
Claude Opus 4.6

Claude Opus 4.6

openrouter/claude-opus-46

opus 4.6 is anthropic’s strongest model for coding and long-running professional tasks. it is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time.

chat
DeepSeek V3.2

DeepSeek V3.2

openrouter/deepseek-v32

deepseek v3.2

chat
Kimi K2

Kimi K2

openrouter/kimi-k2

kimi k2 is the latest, most capable version of open-source model from moonshot ai. built as a thinking agent, scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls.

chat
Kimi K2 Thinking

Kimi K2 Thinking

openrouter/kimi-k2-thinking

a powerful open-source thinking agent that excels at complex, multi-step problem-solving and consistently uses tools effectively over extended operations.

chat
Qwen3 30B A3B

Qwen3 30B A3B

infsh/qwen3-30b-a3b

a powerful language application that excels at multilingual communication and complex task execution, designed for fast performance.

chat
Mistral Small 3.2 24B Instruct

Mistral Small 3.2 24B Instruct

infsh/mistral-small-3-2-24b-it-2506

follows precise instructions, excels at function/tool calling, and can process both text and images for tasks like document understanding and content generation.

chat
Intellect 3

Intellect 3

openrouter/intellect-3

intellect 3

chat
Claude Opus 4.5

Claude Opus 4.5

openrouter/claude-opus-45

claude opus 4.5

chat
Gemini 3 Pro Preview

Gemini 3 Pro Preview

openrouter/gemini-3-pro-preview

gemini 3 pro preview

chat
Claude Sonnet 4.5

Claude Sonnet 4.5

openrouter/claude-sonnet-45

claude sonnet 4.5 is anthropic’s most advanced sonnet model to date, optimized for real-world agents and coding workflows. it delivers state-of-the-art performance on coding benchmarks such as swe-bench verified, with improvements across system design, code security, and specification adherence. the model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.

chat
Claude Haiku 4.5

Claude Haiku 4.5

openrouter/claude-haiku-45

a very fast and economical ai designed for real-time uses and everyday business and coding tasks, offering performance similar to much larger, pricier options.

chat
Magistral Small 2506

Magistral Small 2506

infsh/magistral-small-2506

a high-performance system for complex reasoning, coding, and math problems, featuring strong multilingual support.

chat
GLM-4.6

GLM-4.6

openrouter/glm-46

a powerful, open-source language system excelling in advanced coding, complex reasoning, and integrating tools for sophisticated tasks.

chat

explore more on inference shell

reasoning is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.

we use cookies

we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.

by clicking "accept", you agree to our use of cookies.
learn more.