open weights
24 open weights models and apps on inference shell.
run via API, SDK, or belt CLI.
decision
11
d1-omni-600M
liquid/d1-omni-600m
liquid ai's 587m decision model for text, images and speech. send a state with images or an audio clip and typed questions (choice, score, yes/no); get a probability for every answer. research release

d1-3B
liquid/d1-3b
liquid ai's open 3b decision model for text and images. send a state and up to 8 images with typed questions (choice, score, yes/no); get a probability for every answer in one pass, no generation.

Perplexity Decider v1.1 27B
perplexity/decider-v1-1-27b
perplexity's open decision model for text and images. send a state, up to 4 images and typed questions (choice, score, yes/no); get a calibrated probability for every answer, with no text generation.

Decision 2.0 Vega 27B
vllm-sr/decision-2-0-vega-27b
the largest decision 2.0 model from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one pass. 29.37b parameters, 32k-token input.

JEV-27B-VL
autotrust/jev-27b-vl
autotrust's open decision model with vision. send text and up to 8 images with typed questions (choice, score, yes/no); get a calibrated probability for every answer, with no text generation.

Decision 2.0 Lux 9B
infsh/decision-2-0-lux-9b
the 9b decision 2.0 model from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 7.94b parameters, 16,384-token input, about 18.4 ms per question.

Decision 2.0 Nox 4B
infsh/decision-2-0-nox-4b
the 4b decision 2.0 model from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 4.21b parameters, 16,384-token input, about 12.9 ms per question.

Decision 2.0 Sol 2B Reasoning
infsh/decision-2-0-sol-2b-reasoning
the reasoning variant of decision 2.0 sol 2b, stronger on multi-step problems (arithmetic, code execution, causal and logical questions) at the same speed from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 1.88b parameters, 16,384-token input, about 7.5 ms per question.

Decision 2.0 Sol 2B
infsh/decision-2-0-sol-2b
the 2b decision 2.0 model from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 1.88b parameters, 16,384-token input, about 7.2 ms per question.

Decision 2.0 Eos 0.8B
infsh/decision-2-0-eos-0-8b
a small, fast decision 2.0 model from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 0.75b parameters, 16,384-token input, about 6.0 ms per question.

Decision 2.0 Kai 0.6B
infsh/decision-2-0-kai-0-6b
the smallest and fastest decision 2.0 model from vllm semantic router. send a state and typed questions (choice, score, yes/no); get a probability for every answer in one forward pass, with no text generation. 0.60b parameters, 8,192-token input, about 4.9 ms per question.
chat
10
Kimi K3
melious/kimi-k3
kimi k3 — moonshot's 2.8t open-weight multimodal agentic model with native vision and 1m context for long-horizon coding and reasoning. via the melious api.

GLM 5.2
melious/glm-5-2
glm 5.2 — zai's 744b moe with 1m context and hybrid reasoning for coding and agentic tasks. mit license. via the melious api.

GLM 5.2
openrouter/glm-5-2
glm 5.2 - zhipu's latest flagship language model with 1m context via openrouter

Kimi K2
openrouter/kimi-k2
kimi k2 is the latest, most capable version of open-source model from moonshot ai. built as a thinking agent, scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls.

Kimi K2 Thinking
openrouter/kimi-k2-thinking
a powerful open-source thinking agent that excels at complex, multi-step problem-solving and consistently uses tools effectively over extended operations.

Qwen3 30B A3B
infsh/qwen3-30b-a3b
a powerful language application that excels at multilingual communication and complex task execution, designed for fast performance.

Mistral Small 3.2 24B Instruct
infsh/mistral-small-3-2-24b-it-2506
follows precise instructions, excels at function/tool calling, and can process both text and images for tasks like document understanding and content generation.

Gemma 3n E4B Instruct
infsh/gemma-3n-e4b-it
a fast and versatile tool that can analyze and respond to information from text, images, and audio, designed to run efficiently on small or limited devices.

Magistral Small 2506
infsh/magistral-small-2506
a high-performance system for complex reasoning, coding, and math problems, featuring strong multilingual support.

GLM-4.6
openrouter/glm-46
a powerful, open-source language system excelling in advanced coding, complex reasoning, and integrating tools for sophisticated tasks.
explore more on inference shell
open weights is one of many things you can run on the grid. discover hundreds of apps across image, video, audio, and more.
we use cookies
we use cookies to ensure you get the best experience on our website. for more information on how we use cookies, please see our cookie policy.
by clicking "accept", you agree to our use of cookies.
learn more.
![FLUX.2 [klein]](https://cloud.inference.sh/app/files/t/65hmp52a/syvu1i0n.png)

