# chat apps

> language models from every major lab. 60 apps on inference shell.

**URL:** https://inference.sh/apps/category/chat

---

- [gemini-3-7-flash](https://inference.sh/apps/google/gemini-3-7-flash.md) `google/gemini-3-7-flash`: $1.50/M input, $7.50/M output
- [gemini-3-8-flash](https://inference.sh/apps/google/gemini-3-8-flash.md) `google/gemini-3-8-flash`: $1.50/M input, $7.50/M output
- [DeepSeek V4.1 Flash](https://inference.sh/apps/melious/deepseek-v4-1-flash.md) `melious/deepseek-v4-1-flash`: $0.2273/M input, $0.0114/M cached input, $1.1367/M output
- [MiniMax M3](https://inference.sh/apps/melious/minimax-m3.md) `melious/minimax-m3`: $0.4547/M input, $0.1137/M cached input, $2.2734/M output
- [GLM 5.3](https://inference.sh/apps/melious/glm-5-3.md) `melious/glm-5-3`: $1.1367/M input, $0.2273/M cached input, $3.4101/M output
- [Kimi K3](https://inference.sh/apps/melious/kimi-k3.md) `melious/kimi-k3`: $3.1828/M input, $0.7957/M cached input, $15.9138/M output
- [Qwen 3.6 27B](https://inference.sh/apps/melious/qwen3-6-27b.md) `melious/qwen3-6-27b`: $0.2273/M input, $1.364/M output
- [GLM 5.3 Flash](https://inference.sh/apps/melious/glm-5-3-flash.md) `melious/glm-5-3-flash`: $0.1137/M input, $0.0227/M cached input, $0.4547/M output
- [DeepSeek V4 Pro 0813](https://inference.sh/apps/melious/deepseek-v4-pro-0813.md) `melious/deepseek-v4-pro-0813`: $1.1367/M input, $0.1137/M cached input, $3.4101/M output
- [GLM 5.2](https://inference.sh/apps/melious/glm-5-2.md) `melious/glm-5-2`: $1.1367/M input, $0.2842/M cached input, $4.5468/M output
- [DeepSeek V4 Flash 0731](https://inference.sh/apps/melious/deepseek-v4-flash-0731.md) `melious/deepseek-v4-flash-0731`: $0.1137/M input, $0.0227/M cached input, $0.2842/M output
- [grok-4-7](https://inference.sh/apps/xai/grok-4-7.md) `xai/grok-4-7`: $2.00/M input, $0.50/M cached, $6.00/M output (2x at 200k+ prompt tokens)
- [grok-build-0-1](https://inference.sh/apps/xai/grok-build-0-1.md) `xai/grok-build-0-1`: $1.00/M input, $0.20/M cached, $2.00/M output (2x at 200k+ prompt tokens)
- [grok-4-3](https://inference.sh/apps/xai/grok-4-3.md) `xai/grok-4-3`: $1.25/M input, $0.20/M cached, $2.50/M output (2x at 200k+ prompt tokens)
- [GPT-5.6 Sol](https://inference.sh/apps/openai/gpt-5-6-sol.md) `openai/gpt-5-6-sol`: $4.00/M input, $0.40/M cached input, $20.00/M output
- [GPT-5.6 Luna](https://inference.sh/apps/openai/gpt-5-6-luna.md) `openai/gpt-5-6-luna`: $0.20/M input, $0.02/M cached input, $1.20/M output
- [GPT-5.6 Terra](https://inference.sh/apps/openai/gpt-5-6-terra.md) `openai/gpt-5-6-terra`: $2.00/M input, $0.20/M cached input, $12.00/M output
- [GPT-6 Astra](https://inference.sh/apps/openai/gpt-6-astra.md) `openai/gpt-6-astra`: $10.00/M input, $1.00/M cached, $50.00/M output
- [Gemini 2.5 Flash-Lite](https://inference.sh/apps/google/gemini-2-5-flash-lite.md) `google/gemini-2-5-flash-lite`: $0.10/M input, $0.40/M output
- [Gemini 2.5 Flash](https://inference.sh/apps/google/gemini-2-5-flash.md) `google/gemini-2-5-flash`: $0.30/M input, $2.50/M output
- [Gemini 2.5 Pro](https://inference.sh/apps/google/gemini-2-5-pro.md) `google/gemini-2-5-pro`: $1.25/M input, $10.00/M output
- [MiniMax M2.7](https://inference.sh/apps/minimax/m-2-7.md) `minimax/m-2-7`: $0.30/M input, $1.20/M output
- [MiniMax M3](https://inference.sh/apps/minimax/m3.md) `minimax/m3`: $0.30/M input, $1.20/M output
- [MiniMax M2.7 Highspeed](https://inference.sh/apps/minimax/m-2-7-highspeed.md) `minimax/m-2-7-highspeed`: $0.30/M input, $1.20/M output
- [Claude Opus 5](https://inference.sh/apps/anthropic/claude-opus-5.md) `anthropic/claude-opus-5`: $5.00/M input, $25.00/M output
- [Gemini 3.5 Flash-Lite](https://inference.sh/apps/google/gemini-3-5-flash-lite.md) `google/gemini-3-5-flash-lite`: $0.30/M input, $2.50/M output
- [Gemini 3.6 Flash](https://inference.sh/apps/google/gemini-3-6-flash.md) `google/gemini-3-6-flash`: $1.50/M input, $7.50/M output
- [Claude Sonnet 5](https://inference.sh/apps/anthropic/claude-sonnet-5.md) `anthropic/claude-sonnet-5`: $2.00/M input, $10.00/M output
- [Claude Opus 4.8](https://inference.sh/apps/anthropic/claude-opus-4-8.md) `anthropic/claude-opus-4-8`: $5.00/M input, $25.00/M output
- [GLM 5.2](https://inference.sh/apps/openrouter/glm-5-2.md) `openrouter/glm-5-2`: $1.00/M input, $4.00/M output
- [Claude Mythos 5](https://inference.sh/apps/anthropic/claude-mythos-5.md) `anthropic/claude-mythos-5`: $10.00/M input, $50.00/M output, $1.00/M cache read, $12.50/M cache write
- [Claude Fable 5](https://inference.sh/apps/anthropic/claude-fable-5.md) `anthropic/claude-fable-5`: $10.00/M input, $50.00/M output
- [Qwen3 32B](https://inference.sh/apps/openrouter/qwen3-32b.md) `openrouter/qwen3-32b`: $0.08/M input, $0.28/M output
- [Qwen3 8B](https://inference.sh/apps/openrouter/qwen3-8b.md) `openrouter/qwen3-8b`: $0.05/M input, $0.40/M output
- [Claude Haiku 4.5](https://inference.sh/apps/anthropic/claude-haiku-4-5.md) `anthropic/claude-haiku-4-5`: $1.00/M input, $5.00/M output
- [Claude Sonnet 4.5](https://inference.sh/apps/anthropic/claude-sonnet-4-5.md) `anthropic/claude-sonnet-4-5`: $3.00/M input, $15.00/M output
- [Claude Sonnet 4.6](https://inference.sh/apps/anthropic/claude-sonnet-4-6.md) `anthropic/claude-sonnet-4-6`: $3.00/M input, $15.00/M output, $0.30/M cache read, $3.75/M cache write
- [Claude Opus 4.6](https://inference.sh/apps/anthropic/claude-opus-4-6.md) `anthropic/claude-opus-4-6`: $5.00/M input, $25.00/M output
- [Claude Opus 4.7](https://inference.sh/apps/anthropic/claude-opus-4-7.md) `anthropic/claude-opus-4-7`: $5.00/M input, $25.00/M output
- [Claude Sonnet 4.6](https://inference.sh/apps/openrouter/claude-sonnet-46.md) `openrouter/claude-sonnet-46`: $3.00/M input, $15.00/M output
- [Hunyuan Turbos Hy3 Preview](https://inference.sh/apps/openrouter/hy3-preview.md) `openrouter/hy3-preview`: $0.066/M input, $0.26/M output
- [Gemini 3 Flash Preview](https://inference.sh/apps/openrouter/gemini-3-flash-preview.md) `openrouter/gemini-3-flash-preview`: $0.50/M input, $3.00/M output
- [Kimi K2.6](https://inference.sh/apps/openrouter/kimi-k26.md) `openrouter/kimi-k26`: $0.75/M input, $3.50/M output
- [Claude Opus 4.7](https://inference.sh/apps/openrouter/claude-opus-47.md) `openrouter/claude-opus-47`: $5.00/M input + $25.00/M output tokens
- [Kimi K2.5](https://inference.sh/apps/openrouter/kimi-k25.md) `openrouter/kimi-k25`: $0.4/M input tokens, $1.9/M output tokens
- [GPT-OSS Safeguard 20B](https://inference.sh/apps/openrouter/gpt-oss-safeguard-20b.md) `openrouter/gpt-oss-safeguard-20b`: $0.075/M input tokens, $0.30/M output tokens
- [Claude Opus 4.6](https://inference.sh/apps/openrouter/claude-opus-46.md) `openrouter/claude-opus-46`: $5.00/M input, $25.00/M output
- [DeepSeek V3.2](https://inference.sh/apps/openrouter/deepseek-v32.md) `openrouter/deepseek-v32`: $0.252/M input tokens, $0.378/M output tokens
- [Kimi K2](https://inference.sh/apps/openrouter/kimi-k2.md) `openrouter/kimi-k2`: $0.57/M input tokens, $2.30/M output tokens
- [Kimi K2 Thinking](https://inference.sh/apps/openrouter/kimi-k2-thinking.md) `openrouter/kimi-k2-thinking`: $0.60/M input tokens, $2.50/M output tokens
- [Qwen3 30B A3B](https://inference.sh/apps/infsh/qwen3-30b-a3b.md) `infsh/qwen3-30b-a3b`: GPU time, billed per millisecond
- [Mistral Small 3.2 24B Instruct](https://inference.sh/apps/infsh/mistral-small-3-2-24b-it-2506.md) `infsh/mistral-small-3-2-24b-it-2506`: GPU time, billed per millisecond
- [Intellect 3](https://inference.sh/apps/openrouter/intellect-3.md) `openrouter/intellect-3`: $0.20/M input tokens, $1.10/M output tokens
- [Claude Opus 4.5](https://inference.sh/apps/openrouter/claude-opus-45.md) `openrouter/claude-opus-45`: Priced at $5/M input tokens, $25/M output tokens
- [Gemini 3 Pro Preview](https://inference.sh/apps/openrouter/gemini-3-pro-preview.md) `openrouter/gemini-3-pro-preview`: $2.00/M input tokens, $12.00/M output tokens
- [Gemma 3n E4B Instruct](https://inference.sh/apps/infsh/gemma-3n-e4b-it.md) `infsh/gemma-3n-e4b-it`: GPU time, billed per millisecond
- [Claude Sonnet 4.5](https://inference.sh/apps/openrouter/claude-sonnet-45.md) `openrouter/claude-sonnet-45`: Priced at $3/M input tokens, $15/M output tokens
- [Claude Haiku 4.5](https://inference.sh/apps/openrouter/claude-haiku-45.md) `openrouter/claude-haiku-45`: $1.00/M input tokens, $5.00/M output tokens (minimum charge: $0.002)
- [Magistral Small 2506](https://inference.sh/apps/infsh/magistral-small-2506.md) `infsh/magistral-small-2506`: GPU time, billed per millisecond
- [GLM-4.6](https://inference.sh/apps/openrouter/glm-46.md) `openrouter/glm-46`: $0.43/M input tokens, $1.74/M output tokens