chat
language models from every major lab. 60 apps on inference shell.
run via API, SDK, or belt CLI.
vision
33
gemini-3-7-flash
google/gemini-3-7-flash
gemini 3.7 flash via vertex ai — improved software engineering, web development and agentic workflows over 3.6 flash. 1m token context window.

gemini-3-8-flash
google/gemini-3-8-flash
gemini 3.8 flash via vertex ai — google's most capable flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows. 1m token context window.

DeepSeek V4.1 Flash
melious/deepseek-v4-1-flash
deepseek v4.1 flash — 552b multimodal moe with 1m context, image input and controllable reasoning. frontier agentic and coding performance. via the melious api.

MiniMax M3
melious/minimax-m3
minimax m3 — 428b multimodal moe with 1m context and image input for long-horizon agentic, coding and cowork tasks. via the melious api.

Kimi K3
melious/kimi-k3
kimi k3 — moonshot's 2.8t open-weight multimodal agentic model with native vision and 1m context for long-horizon coding and reasoning. via the melious api.

Qwen 3.6 27B
melious/qwen3-6-27b
qwen 3.6 27b — dense 27b vision-language model with 262k context and thinking mode. strong coding and multimodal reasoning, apache 2.0. via the melious api.

GLM 5.3 Flash
melious/glm-5-3-flash
glm 5.3 flash — zai's 320b moe (18b active), the first natively multimodal glm-5 model. 1m context, image input, low cost. via the melious api.

grok-4-7
xai/grok-4-7
grok 4.7 — xai's most capable model for chat, code and agents. 500k context, reasoning effort low to xhigh, image input and tool use via the xai responses api. direct api.

grok-build-0-1
xai/grok-build-0-1
grok build 0.1 — xai's agentic coding model. 256k context, reasoning, image input and tool use via the xai responses api. direct api.

grok-4-3
xai/grok-4-3
grok 4.3 — xai's fast, low-cost model. 1m context, reasoning effort from none to xhigh, image input and tool use via the xai responses api. direct api.

GPT-5.6 Sol
openai/gpt-5-6-sol
gpt-5.6 sol — openai's flagship gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

GPT-5.6 Luna
openai/gpt-5-6-luna
gpt-5.6 luna — openai's fast, low-cost gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

GPT-5.6 Terra
openai/gpt-5-6-terra
gpt-5.6 terra — openai's balanced gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

GPT-6 Astra
openai/gpt-6-astra
gpt-6 astra — openai's frontier model. 1m context, 128k output, adaptive reasoning (low to max), vision, file input and tool use via the responses api. direct api.

Gemini 2.5 Flash-Lite
google/gemini-2-5-flash-lite
gemini 2.5 flash-lite via vertex ai — most cost-efficient gemini 2.5 model optimized for high-throughput, low-latency workloads. 1m token context window.

Gemini 2.5 Flash
google/gemini-2-5-flash
gemini 2.5 flash via vertex ai — fast, cost-efficient thinking model with strong reasoning, coding, and multimodal performance. 1m token context window.

Gemini 2.5 Pro
google/gemini-2-5-pro
gemini 2.5 pro via vertex ai — google's most capable thinking model with enhanced reasoning, coding, math, and science performance. 1m token context window.

MiniMax M2.7
minimax/m-2-7
minimax-m2.7 — large language model with enhanced reasoning, image understanding, and file processing. 200k context. direct minimax api.

MiniMax M3
minimax/m3
minimax-m3 — frontier multimodal model with 1m context window. text, image, and video inputs. advanced coding, reasoning, and long-horizon agentic tasks. direct minimax api.

Claude Opus 5
anthropic/claude-opus-5
claude opus 5 — anthropic's frontier opus for complex agentic coding and long-horizon work. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

Gemini 3.5 Flash-Lite
google/gemini-3-5-flash-lite
gemini 3.5 flash-lite via vertex ai — fastest, most cost-effective 3.5-class model at 350 output tokens/s. built for high-throughput agentic workflows.

Gemini 3.6 Flash
google/gemini-3-6-flash
gemini 3.6 flash via vertex ai — efficient workhorse model with improved coding, knowledge work, multimodal performance, and 17% fewer output tokens than 3.5 flash.

Claude Sonnet 5
anthropic/claude-sonnet-5
claude sonnet 5 — frontier sonnet with near-opus performance. 1m context, vision, extended thinking, tool use. direct api.

Claude Opus 4.8
anthropic/claude-opus-4-8
claude opus 4.8 — anthropic's most capable opus model. 1m context, 128k output, vision, extended thinking, tool use. direct api.

Claude Mythos 5
anthropic/claude-mythos-5
claude mythos 5 — project glasswing. successor to claude mythos preview. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

Claude Fable 5
anthropic/claude-fable-5
claude fable 5 — anthropic's most capable widely released model. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

Claude Haiku 4.5
anthropic/claude-haiku-4-5
claude haiku 4.5 — fastest and most affordable claude. 200k context, 64k output, vision, extended thinking, tool use. direct api.

Claude Sonnet 4.5
anthropic/claude-sonnet-4-5
claude sonnet 4.5 — previous generation sonnet. 200k context, 64k output, vision, extended thinking, tool use. direct api.

Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
claude sonnet 4.6 — best balance of speed and intelligence. 1m context, 64k output, vision, extended thinking, tool use. direct api.

Claude Opus 4.6
anthropic/claude-opus-4-6
claude opus 4.6 — previous generation opus. 1m context, 128k output, vision, extended thinking, tool use. direct api.

Claude Opus 4.7
anthropic/claude-opus-4-7
claude opus 4.7 — anthropic's most capable model. 1m context, 128k output, vision, extended thinking, tool use. direct api.

Mistral Small 3.2 24B Instruct
infsh/mistral-small-3-2-24b-it-2506
follows precise instructions, excels at function/tool calling, and can process both text and images for tasks like document understanding and content generation.

Gemma 3n E4B Instruct
infsh/gemma-3n-e4b-it
a fast and versatile tool that can analyze and respond to information from text, images, and audio, designed to run efficiently on small or limited devices.

GLM 5.3
melious/glm-5-3
glm 5.3 — zai's 744b moe post-trained for complex coding and long-horizon agentic tasks. 1m context, thinking on by default. via the melious api.

DeepSeek V4 Pro 0813
melious/deepseek-v4-pro-0813
deepseek v4 pro 0813 — 1.6t moe with 1m context and low/high/max reasoning effort for agentic and coding work. text-only. via the melious api.

GLM 5.2
melious/glm-5-2
glm 5.2 — zai's 744b moe with 1m context and hybrid reasoning for coding and agentic tasks. mit license. via the melious api.

DeepSeek V4 Flash 0731
melious/deepseek-v4-flash-0731
deepseek v4 flash 0731 — 304b moe with 1m context and low/high/max reasoning effort. strong agentic coding and tool use. text-only. via the melious api.

MiniMax M2.7 Highspeed
minimax/m-2-7-highspeed
minimax-m2.7-highspeed — m2.7 performance with significantly accelerated inference. 200k context. direct minimax api.

GLM 5.2
openrouter/glm-5-2
glm 5.2 - zhipu's latest flagship language model with 1m context via openrouter

Qwen3 32B
openrouter/qwen3-32b
qwen3 32b - powerful dense language model with reasoning and tool use capabilities via openrouter

Qwen3 8B
openrouter/qwen3-8b
qwen3 8b - efficient dense language model with reasoning and tool use capabilities via openrouter

Claude Sonnet 4.6
openrouter/claude-sonnet-46
sonnet 4.

Hunyuan Turbos Hy3 Preview
openrouter/hy3-preview
hy3 preview is a high-efficiency mixture-of-experts model from tencent designed for agentic workflows and production use.

Gemini 3 Flash Preview
openrouter/gemini-3-flash-preview
gemini 3 flash preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

Kimi K2.6
openrouter/kimi-k26
kimi k2.

Claude Opus 4.7
openrouter/claude-opus-47
opus 4.

Kimi K2.5
openrouter/kimi-k25
kimi k2.5 is the latest model from moonshot ai, building on k2 with enhanced reasoning and capabilities.

GPT-OSS Safeguard 20B
openrouter/gpt-oss-safeguard-20b
gpt-oss safeguard 20b

Claude Opus 4.6
openrouter/claude-opus-46
opus 4.6 is anthropic’s strongest model for coding and long-running professional tasks. it is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time.

DeepSeek V3.2
openrouter/deepseek-v32
deepseek v3.2

Kimi K2
openrouter/kimi-k2
kimi k2 is the latest, most capable version of open-source model from moonshot ai. built as a thinking agent, scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls.

Kimi K2 Thinking
openrouter/kimi-k2-thinking
a powerful open-source thinking agent that excels at complex, multi-step problem-solving and consistently uses tools effectively over extended operations.

Qwen3 30B A3B
infsh/qwen3-30b-a3b
a powerful language application that excels at multilingual communication and complex task execution, designed for fast performance.

Intellect 3
openrouter/intellect-3
intellect 3

Claude Opus 4.5
openrouter/claude-opus-45
claude opus 4.5

Gemini 3 Pro Preview
openrouter/gemini-3-pro-preview
gemini 3 pro preview

Claude Sonnet 4.5
openrouter/claude-sonnet-45
claude sonnet 4.5 is anthropic’s most advanced sonnet model to date, optimized for real-world agents and coding workflows. it delivers state-of-the-art performance on coding benchmarks such as swe-bench verified, with improvements across system design, code security, and specification adherence. the model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.

Claude Haiku 4.5
openrouter/claude-haiku-45
a very fast and economical ai designed for real-time uses and everyday business and coding tasks, offering performance similar to much larger, pricier options.

Magistral Small 2506
infsh/magistral-small-2506
a high-performance system for complex reasoning, coding, and math problems, featuring strong multilingual support.

GLM-4.6
openrouter/glm-46
a powerful, open-source language system excelling in advanced coding, complex reasoning, and integrating tools for sophisticated tasks.
explore more on inference shell
chat is one of the categories on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.