vision
33 vision models and apps on inference shell.
run via API, SDK, or belt CLI.
chat
28
DeepSeek V4.1 Flash
melious/deepseek-v4-1-flash
deepseek v4.1 flash — 552b multimodal moe with 1m context, image input and controllable reasoning. frontier agentic and coding performance. via the melious api.

MiniMax M3
melious/minimax-m3
minimax m3 — 428b multimodal moe with 1m context and image input for long-horizon agentic, coding and cowork tasks. via the melious api.

Kimi K3
melious/kimi-k3
kimi k3 — moonshot's 2.8t open-weight multimodal agentic model with native vision and 1m context for long-horizon coding and reasoning. via the melious api.

Qwen 3.6 27B
melious/qwen3-6-27b
qwen 3.6 27b — dense 27b vision-language model with 262k context and thinking mode. strong coding and multimodal reasoning, apache 2.0. via the melious api.

GLM 5.3 Flash
melious/glm-5-3-flash
glm 5.3 flash — zai's 320b moe (18b active), the first natively multimodal glm-5 model. 1m context, image input, low cost. via the melious api.

grok-4-7
xai/grok-4-7
grok 4.7 — xai's most capable model for chat, code and agents. 500k context, reasoning effort low to xhigh, image input and tool use via the xai responses api. direct api.

grok-build-0-1
xai/grok-build-0-1
grok build 0.1 — xai's agentic coding model. 256k context, reasoning, image input and tool use via the xai responses api. direct api.

grok-4-3
xai/grok-4-3
grok 4.3 — xai's fast, low-cost model. 1m context, reasoning effort from none to xhigh, image input and tool use via the xai responses api. direct api.

GPT-5.6 Sol
openai/gpt-5-6-sol
gpt-5.6 sol — openai's flagship gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

GPT-5.6 Luna
openai/gpt-5-6-luna
gpt-5.6 luna — openai's fast, low-cost gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

GPT-5.6 Terra
openai/gpt-5-6-terra
gpt-5.6 terra — openai's balanced gpt-5.6 model. 1m context, 128k output, reasoning effort none to max, vision, file input and tool use. direct api.

GPT-6 Astra
openai/gpt-6-astra
gpt-6 astra — openai's frontier model. 1m context, 128k output, adaptive reasoning (low to max), vision, file input and tool use via the responses api. direct api.

Gemini 2.5 Flash
google/gemini-2-5-flash
gemini 2.5 flash via vertex ai — fast, cost-efficient thinking model with strong reasoning, coding, and multimodal performance. 1m token context window.

MiniMax M2.7
minimax/m-2-7
minimax-m2.7 — large language model with enhanced reasoning, image understanding, and file processing. 200k context. direct minimax api.

MiniMax M3
minimax/m3
minimax-m3 — frontier multimodal model with 1m context window. text, image, and video inputs. advanced coding, reasoning, and long-horizon agentic tasks. direct minimax api.

Claude Opus 5
anthropic/claude-opus-5
claude opus 5 — anthropic's frontier opus for complex agentic coding and long-horizon work. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

Gemini 3.6 Flash
google/gemini-3-6-flash
gemini 3.6 flash via vertex ai — efficient workhorse model with improved coding, knowledge work, multimodal performance, and 17% fewer output tokens than 3.5 flash.

Claude Sonnet 5
anthropic/claude-sonnet-5
claude sonnet 5 — frontier sonnet with near-opus performance. 1m context, vision, extended thinking, tool use. direct api.

Claude Opus 4.8
anthropic/claude-opus-4-8
claude opus 4.8 — anthropic's most capable opus model. 1m context, 128k output, vision, extended thinking, tool use. direct api.

Claude Mythos 5
anthropic/claude-mythos-5
claude mythos 5 — project glasswing. successor to claude mythos preview. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

Claude Fable 5
anthropic/claude-fable-5
claude fable 5 — anthropic's most capable widely released model. 1m context, 128k output, adaptive thinking, vision, tool use. direct api.

Claude Haiku 4.5
anthropic/claude-haiku-4-5
claude haiku 4.5 — fastest and most affordable claude. 200k context, 64k output, vision, extended thinking, tool use. direct api.

Claude Sonnet 4.5
anthropic/claude-sonnet-4-5
claude sonnet 4.5 — previous generation sonnet. 200k context, 64k output, vision, extended thinking, tool use. direct api.

Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
claude sonnet 4.6 — best balance of speed and intelligence. 1m context, 64k output, vision, extended thinking, tool use. direct api.

Claude Opus 4.6
anthropic/claude-opus-4-6
claude opus 4.6 — previous generation opus. 1m context, 128k output, vision, extended thinking, tool use. direct api.

Claude Opus 4.7
anthropic/claude-opus-4-7
claude opus 4.7 — anthropic's most capable model. 1m context, 128k output, vision, extended thinking, tool use. direct api.

Mistral Small 3.2 24B Instruct
infsh/mistral-small-3-2-24b-it-2506
follows precise instructions, excels at function/tool calling, and can process both text and images for tasks like document understanding and content generation.

Gemma 3n E4B Instruct
infsh/gemma-3n-e4b-it
a fast and versatile tool that can analyze and respond to information from text, images, and audio, designed to run efficiently on small or limited devices.
decision
5
OpenAI Decisions
openai/decisions
openai's decisions api on gpt-6 luna. send text and images with typed questions (choice, score, yes/no); get a probability for every answer in near real time, with no text generation. direct api.

d1-omni-600M
liquid/d1-omni-600m
liquid ai's 587m decision model for text, images and speech. send a state with images or an audio clip and typed questions (choice, score, yes/no); get a probability for every answer. research release

d1-3B
liquid/d1-3b
liquid ai's open 3b decision model for text and images. send a state and up to 8 images with typed questions (choice, score, yes/no); get a probability for every answer in one pass, no generation.

Perplexity Decider v1.1 27B
perplexity/decider-v1-1-27b
perplexity's open decision model for text and images. send a state, up to 4 images and typed questions (choice, score, yes/no); get a calibrated probability for every answer, with no text generation.

JEV-27B-VL
autotrust/jev-27b-vl
autotrust's open decision model with vision. send text and up to 8 images with typed questions (choice, score, yes/no); get a calibrated probability for every answer, with no text generation.
explore more on inference shell
vision is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.