# inference shell - complete documentation & blog > the ai runtime that compounds with every session. run any model, compose agents, stack knowledge - it never forgets. for the summary version, see: https://inference.sh/llms.txt --- # DOCUMENTATION --- # BLOG --- # APPS run any of these with `belt app run ` --- ## Featured Apps ### bytedance/seedance-2-5 **URL:** https://inference.sh/apps/bytedance/seedance-2-5 **Run:** `belt app run bytedance/seedance-2-5` Professional multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.5 model via BytePlus ARK API. Supports up to 4K (10-bit color), durations up to 30s, MOV output, and multimodal reference-to-video with synchronized audio. --- ### pruna/p-video-avatar **URL:** https://inference.sh/apps/pruna/p-video-avatar **Run:** `belt app run pruna/p-video-avatar` Generate talking head videos from a portrait image with text or audio-driven speech --- ## All Apps ### infsh/dramabox **URL:** https://inference.sh/apps/infsh/dramabox Expressive text-to-speech with voice cloning. The prompt controls emotion, laughs, sighs, pauses and delivery. --- ### openai/gpt-image-2-5-sunburst **URL:** https://inference.sh/apps/openai/gpt-image-2-5-sunburst GPT Image 2.5 Sunburst — OpenAI's most capable image model, built for premium creative and editing workflows with tighter control across edits. Text-to-image, reference-image editing, mask inpainting, transparent backgrounds, quality up to max. --- ### openai/gpt-image-2-5-flare **URL:** https://inference.sh/apps/openai/gpt-image-2-5-flare GPT Image 2.5 Flare — OpenAI's fast, high-quality everyday image model. Higher quality than GPT Image 2 at 50% lower latency. Text-to-image, reference-image editing, mask inpainting, transparent backgrounds, quality up to max. --- ### openai/gpt-5-6-sol **URL:** https://inference.sh/apps/openai/gpt-5-6-sol GPT-5.6 Sol — OpenAI's flagship GPT-5.6 model. 1M context, 128k output, reasoning effort none to max, vision, file input and tool use. Direct API. --- ### openai/gpt-5-6-luna **URL:** https://inference.sh/apps/openai/gpt-5-6-luna GPT-5.6 Luna — OpenAI's fast, low-cost GPT-5.6 model. 1M context, 128k output, reasoning effort none to max, vision, file input and tool use. Direct API. --- ### openai/gpt-5-6-terra **URL:** https://inference.sh/apps/openai/gpt-5-6-terra GPT-5.6 Terra — OpenAI's balanced GPT-5.6 model. 1M context, 128k output, reasoning effort none to max, vision, file input and tool use. Direct API. --- ### openai/gpt-6-astra **URL:** https://inference.sh/apps/openai/gpt-6-astra GPT-6 Astra — OpenAI's frontier model. 1M context, 128k output, adaptive reasoning (low to max), vision, file input and tool use via the Responses API. Direct API. --- ### pruna/p-video-edit **URL:** https://inference.sh/apps/pruna/p-video-edit Edit videos from text prompts with optional reference images. Up to 15s input, standard or draft quality. --- ### falai/minimax-h3-max **URL:** https://inference.sh/apps/falai/minimax-h3-max MiniMax H3 Max — fal's post-trained H3 for stronger prompt adherence and aesthetics. Text-to-video, image-to-video with end frame, and reference-based generation. 480P/768P, 5-15s, ~3s per 5s clip. --- ### google/gemini-2-5-flash-lite **URL:** https://inference.sh/apps/google/gemini-2-5-flash-lite Gemini 2.5 Flash-Lite via Vertex AI — most cost-efficient Gemini 2.5 model optimized for high-throughput, low-latency workloads. 1M token context window. --- ### google/gemini-2-5-flash **URL:** https://inference.sh/apps/google/gemini-2-5-flash Gemini 2.5 Flash via Vertex AI — fast, cost-efficient thinking model with strong reasoning, coding, and multimodal performance. 1M token context window. --- ### google/gemini-2-5-pro **URL:** https://inference.sh/apps/google/gemini-2-5-pro Gemini 2.5 Pro via Vertex AI — Google's most capable thinking model with enhanced reasoning, coding, math, and science performance. 1M token context window. --- ### mirage/avatar-x **URL:** https://inference.sh/apps/mirage/avatar-x Generate an expressive talking-head video from a portrait image or video reference and an audio track using Mirage Avatar X, with natural lip sync, eye contact and micro-expressions. --- ### x/post-analytics **URL:** https://inference.sh/apps/x/post-analytics Get engagement analytics for specific posts. Returns impressions, engagements, and other metrics over a time period. --- ### x/trends **URL:** https://inference.sh/apps/x/trends Get trending topics by location. Use WOEID 1 for worldwide, or a specific location ID. --- ### x/following-list **URL:** https://inference.sh/apps/x/following-list List accounts a user follows. Returns profiles with bios, follower counts, and join dates. --- ### x/followers-list **URL:** https://inference.sh/apps/x/followers-list List followers of a user. Returns profiles with bios, follower counts, and join dates. --- ### x/post-likers **URL:** https://inference.sh/apps/x/post-likers List users who liked a specific post. Returns user profiles with follower counts. --- ### x/post-quotes **URL:** https://inference.sh/apps/x/post-quotes List posts that quoted a specific post. Returns quote tweets with text, author, and engagement metrics. --- ### x/bookmark-remove **URL:** https://inference.sh/apps/x/bookmark-remove Remove a bookmarked post on X.com. --- ### x/bookmark-add **URL:** https://inference.sh/apps/x/bookmark-add Bookmark a post on X.com. --- ### x/bookmarks-list **URL:** https://inference.sh/apps/x/bookmarks-list List your bookmarked posts. Returns saved posts with text, author, and engagement metrics. --- ### x/dm-list **URL:** https://inference.sh/apps/x/dm-list List recent DM events. Returns direct message text, sender, and timestamps. --- ### x/user-mentions **URL:** https://inference.sh/apps/x/user-mentions Get your recent mentions. Returns posts that mention the authenticated user. --- ### x/user-timeline **URL:** https://inference.sh/apps/x/user-timeline Get your home timeline. Returns recent posts from accounts you follow with engagement metrics. --- ### x/user-posts **URL:** https://inference.sh/apps/x/user-posts List posts by a user. Returns text, timestamps, engagement metrics, and article metadata. --- ### x/probe-timeline **URL:** https://inference.sh/apps/x/probe-timeline Probe user posts and timeline for article metadata --- ### x/article-publish **URL:** https://inference.sh/apps/x/article-publish Publish long-form articles on X.com from markdown. Supports headings, lists, blockquotes, code blocks, GFM tables, LaTeX, images, embedded tweets, and inline formatting (bold, italic, strikethrough, links). Optional cover image and draft-only mode. --- ### eval/fid-score **URL:** https://inference.sh/apps/eval/fid-score Frechet Inception Distance — measures distributional similarity between two sets of images --- ### eval/resize **URL:** https://inference.sh/apps/eval/resize Crop and resize images — center crop to square, face-aware crop, or custom resize --- ### eval/clipscore **URL:** https://inference.sh/apps/eval/clipscore CLIPScore — cosine similarity between CLIP text and image embeddings for prompt adherence --- ### eval/arcface **URL:** https://inference.sh/apps/eval/arcface ArcFace identity similarity — cosine similarity between face embeddings of two images --- ### x/post-render **URL:** https://inference.sh/apps/x/post-render Render X/Twitter post cards as PNG images — dark/light theme, profile picture, engagement metrics --- ### eval/pickscore **URL:** https://inference.sh/apps/eval/pickscore PickScore v1 — scores how well an image matches a text prompt (CLIP-H finetuned on Pick-a-Pic) --- ### eval/blur **URL:** https://inference.sh/apps/eval/blur Degrade images with blur, noise, and compression for benchmark testing --- ### eval/charts **URL:** https://inference.sh/apps/eval/charts Research-quality chart generation (bar, grouped bar, scatter) using matplotlib --- ### archive/download **URL:** https://inference.sh/apps/archive/download download a file from an internet archive item --- ### archive/metadata **URL:** https://inference.sh/apps/archive/metadata get item metadata from the internet archive --- ### archive/search **URL:** https://inference.sh/apps/archive/search search the internet archive --- ### archive/wayback **URL:** https://inference.sh/apps/archive/wayback check url availability in the wayback machine --- ### infsh/markdown-to-word **URL:** https://inference.sh/apps/infsh/markdown-to-word Convert Markdown text or files to Word (.docx) documents --- ### bfl/flux-3-video **URL:** https://inference.sh/apps/bfl/flux-3-video FLUX 3 Video by Black Forest Labs — generate, animate, and extend video up to 20s at HD or Full HD with synchronized audio. Supports text-to-video, image-to-video with keyframes, video continuation, and draft mode. --- ### minimax/music-cover **URL:** https://inference.sh/apps/minimax/music-cover MiniMax Music Cover — AI-powered song covers and style transfer. Upload a reference track and generate a cover with new style and optional new lyrics. --- ### minimax/speech-2-8-hd **URL:** https://inference.sh/apps/minimax/speech-2-8-hd MiniMax Speech 2.8 HD — high-quality text-to-speech. 40 languages, 9 emotions, custom voices. Direct MiniMax API. --- ### minimax/music-3-0 **URL:** https://inference.sh/apps/minimax/music-3-0 MiniMax Music 3.0 — AI music generation up to 5 minutes. Supports vocals with lyrics, instrumentals, and style prompts. Direct MiniMax API. --- ### minimax/speech-2-8-turbo **URL:** https://inference.sh/apps/minimax/speech-2-8-turbo MiniMax Speech 2.8 Turbo — fast text-to-speech. 40 languages, 9 emotions, custom voices. Lower latency than HD. Direct MiniMax API. --- ### minimax/m-2-7 **URL:** https://inference.sh/apps/minimax/m-2-7 MiniMax-M2.7 — large language model with enhanced reasoning, image understanding, and file processing. 200K context. Direct MiniMax API. --- ### minimax/m3 **URL:** https://inference.sh/apps/minimax/m3 MiniMax-M3 — frontier multimodal model with 1M context window. Text, image, and video inputs. Advanced coding, reasoning, and long-horizon agentic tasks. Direct MiniMax API. --- ### minimax/m-2-7-highspeed **URL:** https://inference.sh/apps/minimax/m-2-7-highspeed MiniMax-M2.7-highspeed — M2.7 performance with significantly accelerated inference. 200K context. Direct MiniMax API. --- ### minimax/h3 **URL:** https://inference.sh/apps/minimax/h3 MiniMax H3 — multimodal video generation with native audio. Text-to-video, image-to-video with end_image, and reference-based generation. 2K resolution, 5-15s duration, 24fps. Prompt with timelines, audio cues, and negative lists. --- ### krea/krea-2-medium-turbo-train **URL:** https://inference.sh/apps/krea/krea-2-medium-turbo-train Train custom LoRA styles for Krea 2 Medium Turbo generation --- ### krea/krea-2-medium **URL:** https://inference.sh/apps/krea/krea-2-medium Krea 2 Medium — expressive illustrations, ~10s per generation, 1.5K native resolution --- ### krea/krea-2-medium-train **URL:** https://inference.sh/apps/krea/krea-2-medium-train Train custom LoRA styles for Krea 2 Medium generation --- ### krea/krea-2-large-train **URL:** https://inference.sh/apps/krea/krea-2-large-train Train custom LoRA styles for Krea 2 Large generation --- ### krea/krea-2-large **URL:** https://inference.sh/apps/krea/krea-2-large Krea 2 Large — photorealistic generation, ~25s per generation, 2K native resolution --- ### krea/krea-2-medium-turbo **URL:** https://inference.sh/apps/krea/krea-2-medium-turbo Krea 2 Medium Turbo — fastest K2 model, ~3s per generation, 1.5K native resolution --- ### runway/act-two **URL:** https://inference.sh/apps/runway/act-two Runway Act-Two — character performance transfer. Animate a character image or video using a reference performance video with facial expressions, gestures, and body control. 5 credits/second. --- ### runway/gen-4-image-turbo **URL:** https://inference.sh/apps/runway/gen-4-image-turbo Runway Gen-4 Image Turbo — fast text-to-image generation with reference image support. 2 credits per image at any resolution. --- ### runway/gen-4-image **URL:** https://inference.sh/apps/runway/gen-4-image Runway Gen-4 Image — high-quality text-to-image generation with optional reference images and @tag syntax. 5 credits per 720p, 8 credits per 1080p. --- ### runway/aleph-2 **URL:** https://inference.sh/apps/runway/aleph-2 Runway Aleph 2.0 — video-to-video transformation with text and keyframe guidance. Supports aspect ratio targeting and up to 5 reference keyframes. 28 credits/second. --- ### runway/gen-4-turbo **URL:** https://inference.sh/apps/runway/gen-4-turbo Runway Gen-4 Turbo — fast image-to-video generation. Animate images into video with text guidance. 5 credits/second. --- ### runway/gen-4-5 **URL:** https://inference.sh/apps/runway/gen-4-5 Runway Gen-4.5 — high-quality video generation from text or image. 2-10 second duration, multiple aspect ratios. 12 credits/second. --- ### mirage/text-overlays **URL:** https://inference.sh/apps/mirage/text-overlays Render up to 4 static text variants onto one video with per-variant font, size and colour — built for testing ad hooks and headlines against the same footage. --- ### mirage/video-1 **URL:** https://inference.sh/apps/mirage/video-1 Generate an expressive talking-head video from a portrait image and an audio track using Mirage Video 1, with natural lip sync, eye contact and micro-expressions. --- ### mirage/video-captions **URL:** https://inference.sh/apps/mirage/video-captions Burn styled animated captions onto a vertical video using Mirage caption templates, from an upload or an existing Mirage video ID. --- ### anthropic/claude-opus-5 **URL:** https://inference.sh/apps/anthropic/claude-opus-5 Claude Opus 5 — Anthropic's frontier Opus for complex agentic coding and long-horizon work. 1M context, 128k output, adaptive thinking, vision, tool use. Direct API. --- ### pruna/p-image-ideogram **URL:** https://inference.sh/apps/pruna/p-image-ideogram High-quality text-to-image generation with strong typography and prompt understanding, built with Ideogram --- ### pixverse/modify **URL:** https://inference.sh/apps/pixverse/modify PixVerse Modify — edit video content using text prompts. Swap subjects, add elements, or restyle regions with reference images. --- ### pixverse/avatar **URL:** https://inference.sh/apps/pixverse/avatar PixVerse Avatar — generate talking avatar video from a portrait image. Supports audio file or text-to-speech input at up to 1080p. --- ### pixverse/lip-sync **URL:** https://inference.sh/apps/pixverse/lip-sync PixVerse Lip Sync — align speech to mouth movement in video. Supports audio file input or text-to-speech with selectable voices. --- ### pixverse/fusion **URL:** https://inference.sh/apps/pixverse/fusion PixVerse Fusion — generate video from reference images. Compose 1-3 subject/background images into a video scene with text-guided motion. --- ### pixverse/extend **URL:** https://inference.sh/apps/pixverse/extend PixVerse Extend — extend an existing video with AI-generated continuation. Supports 5s or 8s extensions at up to 1080p. --- ### pixverse/c1 **URL:** https://inference.sh/apps/pixverse/c1 PixVerse C1 — advanced video generation model. Text-to-video and image-to-video with up to 1080p resolution and 15s duration. Per-second billing. --- ### pixverse/v6 **URL:** https://inference.sh/apps/pixverse/v6 PixVerse V6 — versatile video generation model. Text-to-video and image-to-video with up to 1080p resolution and 15s duration. Per-second billing. --- ### google/gemini-3-5-flash-lite **URL:** https://inference.sh/apps/google/gemini-3-5-flash-lite Gemini 3.5 Flash-Lite via Vertex AI — fastest, most cost-effective 3.5-class model at 350 output tokens/s. Built for high-throughput agentic workflows. --- ### google/gemini-3-6-flash **URL:** https://inference.sh/apps/google/gemini-3-6-flash Gemini 3.6 Flash via Vertex AI — efficient workhorse model with improved coding, knowledge work, multimodal performance, and 17% fewer output tokens than 3.5 Flash. --- ### fastly/domain-search **URL:** https://inference.sh/apps/fastly/domain-search domain research api — search for available domains and check registration status --- ### infsh/harrier-oss-v1 **URL:** https://inference.sh/apps/infsh/harrier-oss-v1 Multilingual text embedding using Microsoft's Harrier OSS v1 models. Supports retrieval, clustering, semantic similarity, classification, and reranking with state-of-the-art MTEB v2 scores. --- ### ceramic/search **URL:** https://inference.sh/apps/ceramic/search Web search powered by Ceramic AI --- ### bytedance/seedance-2-0-mini **URL:** https://inference.sh/apps/bytedance/seedance-2-0-mini Cost-effective multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.0 Mini model via BytePlus ARK API. ~50% cheaper than Seedance 2.0, supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio. --- ### bytedance/seedream-5-pro **URL:** https://inference.sh/apps/bytedance/seedream-5-pro ByteDance's flagship Seedream 5.0 Pro image model via BytePlus ARK API. Precision creation and editing with pixel-level regional edits, intelligent layer understanding, complex infographic generation, multi-reference blending (up to 10 images), and native text rendering in 14 languages. --- ### topaz/video-upscale **URL:** https://inference.sh/apps/topaz/video-upscale video upscaling and enhancement — proteus family models for precision upscaling, deinterlacing, face recovery, and cgi enhancement --- ### arxiv/search **URL:** https://inference.sh/apps/arxiv/search search arxiv papers by query with field prefixes, boolean operators, category filtering, and sorting --- ### biorxiv/search **URL:** https://inference.sh/apps/biorxiv/search search biorxiv and medrxiv preprints by date range with optional category filtering --- ### chemrxiv/search **URL:** https://inference.sh/apps/chemrxiv/search search chemrxiv preprints via crossref api --- ### arxiv/paper **URL:** https://inference.sh/apps/arxiv/paper get a specific arxiv paper by its id with full metadata including title, authors, abstract, and links --- ### chemrxiv/paper **URL:** https://inference.sh/apps/chemrxiv/paper get a chemrxiv paper by doi via crossref api --- ### biorxiv/paper **URL:** https://inference.sh/apps/biorxiv/paper get a specific biorxiv or medrxiv paper by its doi --- ### topaz/astra **URL:** https://inference.sh/apps/topaz/astra creative video upscaling — ai-guided upscaling with prompt and creativity controls --- ### topaz/starlight **URL:** https://inference.sh/apps/topaz/starlight generative video upscaling — precision, hq, mini, sharp, and fast models --- ### topaz/denoise **URL:** https://inference.sh/apps/topaz/denoise video denoising — nyx family models for noise, compression, and artifact removal --- ### topaz/video-utilities **URL:** https://inference.sh/apps/topaz/video-utilities video utilities — motion deblur, colorization, and sdr to hdr conversion --- ### topaz/frame-interpolation **URL:** https://inference.sh/apps/topaz/frame-interpolation video frame interpolation — slowmo and fps boost with apollo, chronos, and aion models --- ### topaz/proteus **URL:** https://inference.sh/apps/topaz/proteus proteus video upscaling and enhancement — precision upscaling, deinterlacing, face recovery, and cgi enhancement --- ### you/contents **URL:** https://inference.sh/apps/you/contents contents api — fetch clean markdown or html from any url, batch up to 10 pages per request --- ### you/finance-research **URL:** https://inference.sh/apps/you/finance-research finance research api — agentic financial research with filings, macro data, and institutional-grade sources --- ### you/research **URL:** https://inference.sh/apps/you/research deep research api — multi-step web research with source-backed citations and configurable effort levels --- ### you/search **URL:** https://inference.sh/apps/you/search web search api — ground your apps in reliable, web-scale knowledge with contextual snippets --- ### anthropic/claude-sonnet-5 **URL:** https://inference.sh/apps/anthropic/claude-sonnet-5 Claude Sonnet 5 — Frontier Sonnet with near-Opus performance. 1M context, vision, extended thinking, tool use. Direct API. --- ### anthropic/claude-opus-4-8 **URL:** https://inference.sh/apps/anthropic/claude-opus-4-8 Claude Opus 4.8 — Anthropic's most capable Opus model. 1M context, 128k output, vision, extended thinking, tool use. Direct API. --- ### google/gemini-3-1-flash-lite-image **URL:** https://inference.sh/apps/google/gemini-3-1-flash-lite-image Gemini 3.1 Flash Lite Image (NanoBanana 2 Lite) via Vertex AI — ultra-low latency image generation --- ### google/gemini-omni-flash **URL:** https://inference.sh/apps/google/gemini-omni-flash Gemini Omni Flash — text-to-video with synchronized audio, grounded in real-world knowledge --- ### microsoft/mai-image-2-5 **URL:** https://inference.sh/apps/microsoft/mai-image-2-5 MAI Image 2.5 — Microsoft's photorealistic image generation and editing model with fine-grained pixel-level control. --- ### openrouter/glm-5-2 **URL:** https://inference.sh/apps/openrouter/glm-5-2 GLM 5.2 - zhipu's latest flagship language model with 1M context via OpenRouter --- ### reve/remix **URL:** https://inference.sh/apps/reve/remix Reve Remix — Create images from text and 1-6 reference images combined. --- ### reve/edit **URL:** https://inference.sh/apps/reve/edit Reve Edit — Edit images with natural language instructions. Top 3 on LMArena leaderboard. --- ### reve/create **URL:** https://inference.sh/apps/reve/create Reve Create — Generate images from text with best-in-class prompt adherence and text rendering. --- ### pruna/p-video-replace **URL:** https://inference.sh/apps/pruna/p-video-replace Replace characters in videos using reference images. Preserves motion, timing, camera, and scene. --- ### anthropic/claude-mythos-5 **URL:** https://inference.sh/apps/anthropic/claude-mythos-5 Claude Mythos 5 — Project Glasswing. Successor to Claude Mythos Preview. 1M context, 128k output, adaptive thinking, vision, tool use. Direct API. --- ### anthropic/claude-fable-5 **URL:** https://inference.sh/apps/anthropic/claude-fable-5 Claude Fable 5 — Anthropic's most capable widely released model. 1M context, 128k output, adaptive thinking, vision, tool use. Direct API. --- ### elevenlabs/voice-remix **URL:** https://inference.sh/apps/elevenlabs/voice-remix ElevenLabs Voice Remix - Modify voice characteristics like accent, gender, style, pacing --- ### elevenlabs/voice-clone **URL:** https://inference.sh/apps/elevenlabs/voice-clone ElevenLabs Voice Clone - Instantly clone a voice from audio samples --- ### elevenlabs/voice-design **URL:** https://inference.sh/apps/elevenlabs/voice-design ElevenLabs Voice Design - Create custom AI voices from text descriptions --- ### x/post-search **URL:** https://inference.sh/apps/x/post-search Search recent posts on X.com. Use conversation_id to get replies to a tweet, or any X search query. Returns up to 100 posts with text, author, and engagement metrics. --- ### google/gemini-3-pro-image **URL:** https://inference.sh/apps/google/gemini-3-pro-image Gemini 3 Pro Image (NanoBanana Pro) via Vertex AI - Advanced image generation model powered by Google Cloud --- ### google/gemini-3-1-flash-image **URL:** https://inference.sh/apps/google/gemini-3-1-flash-image Gemini 3.1 Flash Image (NanoBanana 2) via Vertex AI - Advanced image generation model powered by Google Cloud --- ### infsh/deepseek-ocr-2 **URL:** https://inference.sh/apps/infsh/deepseek-ocr-2 Next-gen document OCR with improved math, tables, and reading order. Converts images and PDFs to structured markdown. --- ### heygen/create-avatar **URL:** https://inference.sh/apps/heygen/create-avatar Create HeyGen avatars from video footage (digital twin), a photo (photo avatar), or a text prompt (AI-generated). Returns a look ID for use with avatar-video. --- ### openrouter/qwen3-32b **URL:** https://inference.sh/apps/openrouter/qwen3-32b Qwen3 32B - powerful dense language model with reasoning and tool use capabilities via OpenRouter --- ### openrouter/qwen3-8b **URL:** https://inference.sh/apps/openrouter/qwen3-8b Qwen3 8B - efficient dense language model with reasoning and tool use capabilities via OpenRouter --- ### anthropic/claude-haiku-4-5 **URL:** https://inference.sh/apps/anthropic/claude-haiku-4-5 Claude Haiku 4.5 — Fastest and most affordable Claude. 200k context, 64k output, vision, extended thinking, tool use. Direct API. --- ### anthropic/claude-sonnet-4-5 **URL:** https://inference.sh/apps/anthropic/claude-sonnet-4-5 Claude Sonnet 4.5 — Previous generation Sonnet. 200k context, 64k output, vision, extended thinking, tool use. Direct API. --- ### anthropic/claude-sonnet-4-6 **URL:** https://inference.sh/apps/anthropic/claude-sonnet-4-6 Claude Sonnet 4.6 — Best balance of speed and intelligence. 1M context, 64k output, vision, extended thinking, tool use. Direct API. --- ### anthropic/claude-opus-4-6 **URL:** https://inference.sh/apps/anthropic/claude-opus-4-6 Claude Opus 4.6 — Previous generation Opus. 1M context, 128k output, vision, extended thinking, tool use. Direct API. --- ### anthropic/claude-opus-4-7 **URL:** https://inference.sh/apps/anthropic/claude-opus-4-7 Claude Opus 4.7 — Anthropic's most capable model. 1M context, 128k output, vision, extended thinking, tool use. Direct API. --- ### heygen/text-to-speech **URL:** https://inference.sh/apps/heygen/text-to-speech Generate natural speech audio from text using HeyGen's Starfish TTS engine. Supports configurable voice, speed, SSML input, and multiple languages. --- ### heygen/lipsync **URL:** https://inference.sh/apps/heygen/lipsync Re-sync video lip movements to new audio using HeyGen's lipsync technology. Supports speed and precision modes with optional captioning. --- ### heygen/video-translate **URL:** https://inference.sh/apps/heygen/video-translate Translate videos into 30+ languages with voice cloning and lip-sync using HeyGen. Supports speed and precision modes with optional captioning. --- ### heygen/video-agent **URL:** https://inference.sh/apps/heygen/video-agent Generate complete videos from natural language prompts using HeyGen's AI video agent. The agent handles avatar selection, scripting, and production automatically. --- ### heygen/photo-video **URL:** https://inference.sh/apps/heygen/photo-video Animate portrait photos into talking videos using HeyGen. Upload a face image and add speech with configurable voice, motion prompts, and expressiveness. --- ### heygen/avatar-video **URL:** https://inference.sh/apps/heygen/avatar-video Generate talking avatar videos using HeyGen's digital and photo avatars with Avatar IV or V engines, configurable voice, resolution up to 4K, and expressiveness. --- ### veed/subtitles **URL:** https://inference.sh/apps/veed/subtitles Add professional burned-in subtitles to videos with 25+ style presets. Supports 100+ languages with automatic transcription or custom SRT files. --- ### infsh/html-to-video **URL:** https://inference.sh/apps/infsh/html-to-video Render HTML/CSS/JS animations to video — supports GSAP timelines, CSS animations, Web Animations API --- ### klingai/image-v2 **URL:** https://inference.sh/apps/klingai/image-v2 Kling Image V2 (Kolors V2.0) - text-to-image with 2K resolution, multi-image reference, and restyle. Restyle output matches input resolution. --- ### klingai/image-o1 **URL:** https://inference.sh/apps/klingai/image-o1 Kling Image O1 (Kolors Image-O1) - omni image generation with element control. Text-to-image and image-to-image at 1K/2K. $0.028/image. --- ### klingai/image-3o **URL:** https://inference.sh/apps/klingai/image-3o Kling Image 3O (Kolors Image-3O) - most capable image model with native 4K, series-image generation, and element control. $0.028/image (4K $0.056). --- ### klingai/image-v1 **URL:** https://inference.sh/apps/klingai/image-v1 Kling Image V1 (Kolors V1.0) - basic text-to-image and image-to-image generation. Cheapest option at $0.0035/image. --- ### klingai/image-v1-5 **URL:** https://inference.sh/apps/klingai/image-v1-5 Kling Image V1.5 (Kolors V1.5) - text-to-image with subject and face reference for character consistency. Generate images preserving a person's appearance. --- ### klingai/image-v2-1 **URL:** https://inference.sh/apps/klingai/image-v2-1 Kling Image V2.1 (Kolors V2.1) - text-to-image and multi-image reference generation. Combine multiple images for complex compositions. --- ### klingai/image-v3 **URL:** https://inference.sh/apps/klingai/image-v3 Kling Image V3 (Kolors V3.0) - latest image generation model with 1K/2K resolution support. Highest quality text-to-image. --- ### klingai/video-v3 **URL:** https://inference.sh/apps/klingai/video-v3 Kling V3.0 - latest and most capable video generation model. Native 4K output, multi-shot generation, flexible 3-15s duration billed per second, element control, motion control, and synchronized audio. --- ### klingai/video-v2-6 **URL:** https://inference.sh/apps/klingai/video-v2-6 Kling V2.6 video generation with native sound and voice control. Supports text-to-video and image-to-video with start/end frames, synchronized audio generation, and voice-driven character animation. --- ### klingai/video-o1 **URL:** https://inference.sh/apps/klingai/video-o1 Kling Video O1 (Omni) - unified video generation with text, image references, start/end frames, element references, and video references for editing and style transfer. The most capable Kling model. --- ### klingai/video-v2-5 **URL:** https://inference.sh/apps/klingai/video-v2-5 Kling V2.5 Turbo - fast video generation from text and images. Supports start/end frame interpolation in pro mode. Optimized for speed while maintaining high quality at up to 1080p. --- ### klingai/lip-sync **URL:** https://inference.sh/apps/klingai/lip-sync Kling Lip Sync - drive mouth movements in videos using text or audio. Ideal for dubbing, adding speech to silent videos, or replacing dialogue. --- ### klingai/avatar **URL:** https://inference.sh/apps/klingai/avatar Kling Avatar - generate digital human broadcast-style talking head videos from a single face photo. Provide text or audio for the avatar to speak. --- ### klingai/video-to-audio **URL:** https://inference.sh/apps/klingai/video-to-audio Kling Video-to-Audio - add generated sound effects, ambient audio, or music to any video. Works with Kling-generated and user-uploaded videos (3-20s). --- ### inworld/voice-cloning **URL:** https://inference.sh/apps/inworld/voice-cloning Clone a voice from 5-15 seconds of audio using Inworld instant voice cloning. Use the cloned voice ID with any Inworld TTS model. --- ### inworld/voice-design **URL:** https://inference.sh/apps/inworld/voice-design Design a custom voice from a text description using Inworld AI. Describe the voice you want and get up to 3 previews. Publish the one you like to use with any Inworld TTS model. --- ### inworld/text-to-speech-2 **URL:** https://inference.sh/apps/inworld/text-to-speech-2 Inworld TTS-2 - High-quality multilingual text-to-speech with 100+ languages and natural-language steering --- ### inworld/text-to-speech-1-5-max **URL:** https://inference.sh/apps/inworld/text-to-speech-1-5-max Inworld TTS 1.5 Max - Low-latency text-to-speech with 15 languages (<200ms P50) --- ### inworld/speech-to-text **URL:** https://inference.sh/apps/inworld/speech-to-text Inworld Speech to Text - Multi-provider speech transcription with word timestamps --- ### inworld/text-to-speech-1-5-mini **URL:** https://inference.sh/apps/inworld/text-to-speech-1-5-mini Inworld TTS 1.5 Mini - Ultra-low-latency text-to-speech with 15 languages (~120ms P50) --- ### openrouter/claude-sonnet-46 **URL:** https://inference.sh/apps/openrouter/claude-sonnet-46 Sonnet 4. --- ### openrouter/hy3-preview **URL:** https://inference.sh/apps/openrouter/hy3-preview Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. --- ### openrouter/gemini-3-flash-preview **URL:** https://inference.sh/apps/openrouter/gemini-3-flash-preview Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. --- ### openrouter/kimi-k26 **URL:** https://inference.sh/apps/openrouter/kimi-k26 Kimi K2. --- ### openrouter/claude-opus-47 **URL:** https://inference.sh/apps/openrouter/claude-opus-47 Opus 4. --- ### xai/grok-imagine-image-quality **URL:** https://inference.sh/apps/xai/grok-imagine-image-quality Generate and edit high-quality images using xAI's Grok Imagine Quality model. Supports 1K and 2K output resolutions with text-to-image and image editing. --- ### infsh/hyperframes-render **URL:** https://inference.sh/apps/infsh/hyperframes-render Render HeyGen Hyperframes compositions to video — supports clips, GSAP timelines, track layering --- ### bria/rmbg **URL:** https://inference.sh/apps/bria/rmbg Remove the background from an image, producing a transparent cutout. The general-purpose background removal — for product-specific cutouts, use product-cutout instead. Output can be passed to replace-background, blur-background, or any editing app. --- ### bria/increase-resolution **URL:** https://inference.sh/apps/bria/increase-resolution Upscale images 2x or 4x (max 8192x8192) while preserving original content --- ### bria/expand **URL:** https://inference.sh/apps/bria/expand Expand image canvas with AI-generated content matching the original scene --- ### bria/erase **URL:** https://inference.sh/apps/bria/erase Remove objects from images using mask-based inpainting while preserving quality --- ### bria/generate **URL:** https://inference.sh/apps/bria/generate Generate images from text prompts using Bria Fibo --- ### bria/generate-lite **URL:** https://inference.sh/apps/bria/generate-lite Fast image generation from text prompts using Bria Fibo Lite --- ### bria/structured-prompt **URL:** https://inference.sh/apps/bria/structured-prompt Generate structured prompt JSON from text or images using Bria --- ### bria/ads-generate **URL:** https://inference.sh/apps/bria/ads-generate Generate multiple ads in various sizes from templates and brand assets --- ### bria/product-cutout **URL:** https://inference.sh/apps/bria/product-cutout Cut out product from image with transparent background --- ### bria/gen-fill **URL:** https://inference.sh/apps/bria/gen-fill Generative fill — replace masked regions with AI-generated content guided by a text prompt --- ### bria/product-packshot **URL:** https://inference.sh/apps/bria/product-packshot Generate professional 2000x2000 product packshot images --- ### bria/replace-background **URL:** https://inference.sh/apps/bria/replace-background Replace image background with AI-generated content from a text prompt or reference image --- ### bria/video-rmbg **URL:** https://inference.sh/apps/bria/video-rmbg Remove background from videos with optional color replacement --- ### bria/product-shadow **URL:** https://inference.sh/apps/bria/product-shadow Add realistic shadows to product cutout images --- ### bria/video-eraser **URL:** https://inference.sh/apps/bria/video-eraser Erase objects from video using a mask with inpainting --- ### bria/video-replace-background **URL:** https://inference.sh/apps/bria/video-replace-background Replace video background with an image or another video --- ### bria/edit **URL:** https://inference.sh/apps/bria/edit Edit an image using natural language text instructions --- ### bria/video-increase-resolution **URL:** https://inference.sh/apps/bria/video-increase-resolution Upscale video resolution up to 8K using AI super-resolution --- ### bria/video-green-screen **URL:** https://inference.sh/apps/bria/video-green-screen Apply green or blue screen effect to video foreground --- ### bytedance/seedance-2-0 **URL:** https://inference.sh/apps/bytedance/seedance-2-0 Professional multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.0 model via BytePlus ARK API. Supports up to 4K (10-bit color), text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio. --- ### bytedance/seedance-2-0-fast **URL:** https://inference.sh/apps/bytedance/seedance-2-0-fast Fast multimodal video generation from text, images, video, and audio references using ByteDance's Seedance 2.0 Fast model via BytePlus ARK API. Supports text-to-video, image-to-video, and multimodal reference-to-video with synchronized audio. --- ### alibaba/happyhorse-1-0-video-edit **URL:** https://inference.sh/apps/alibaba/happyhorse-1-0-video-edit HappyHorse 1.0 Video Edit supports advanced video editing through natural language instructions with up to 5 reference images, preserving original motion dynamics via DashScope API --- ### alibaba/happyhorse-1-0-t2v **URL:** https://inference.sh/apps/alibaba/happyhorse-1-0-t2v HappyHorse 1.0 Text-to-Video generates physically realistic videos with smooth motion from text prompts via DashScope API, supporting 720P/1080P resolution and up to 15 seconds duration --- ### alibaba/happyhorse-1-0-i2v **URL:** https://inference.sh/apps/alibaba/happyhorse-1-0-i2v HappyHorse 1.0 Image-to-Video generates physically realistic videos with smooth motion from a single image and optional text description via DashScope API, supporting 720P/1080P resolution --- ### alibaba/happyhorse-1-0-r2v **URL:** https://inference.sh/apps/alibaba/happyhorse-1-0-r2v HappyHorse 1.0 Reference-to-Video generates videos preserving subject characters from up to 9 reference images, with enhanced stability in subject and scene referencing via DashScope API --- ### openai/gpt-image-2 **URL:** https://inference.sh/apps/openai/gpt-image-2 Generate and edit images using OpenAI's GPT Image 2 model. Supports text-to-image, image editing with reference images, mask-based inpainting, and transparent backgrounds. --- ### patina/extract-material **URL:** https://inference.sh/apps/patina/extract-material Extracts a seamlessly tiling texture plus PBR material maps from a region of a source image described by a text prompt, via fal.ai PATINA. --- ### patina/text-to-material **URL:** https://inference.sh/apps/patina/text-to-material Generates seamlessly tiling PBR materials up to 8K from a text prompt (optional image-to-image and inpainting) via fal.ai PATINA. --- ### patina/image-to-material **URL:** https://inference.sh/apps/patina/image-to-material Predicts seamless high-resolution PBR material maps (basecolor, normal, roughness, metalness, height) from a single input image via fal.ai PATINA. --- ### infsh/omnivoice **URL:** https://inference.sh/apps/infsh/omnivoice Zero-shot text-to-speech with voice cloning for 600+ languages. --- ### alibaba/wan-2-7-videoedit **URL:** https://inference.sh/apps/alibaba/wan-2-7-videoedit Wan 2.7 Video Edit performs instruction-based video editing and style transfer using multimodal inputs (text, images, video) via DashScope API with 720P/1080P output --- ### alibaba/wan-2-7-r2v **URL:** https://inference.sh/apps/alibaba/wan-2-7-r2v Wan 2.7 Reference-to-Video generates videos featuring characters from reference images and videos, supporting multi-character interaction, voice timbre cloning, and first-frame control --- ### alibaba/wan-2-7-i2v **URL:** https://inference.sh/apps/alibaba/wan-2-7-i2v Wan 2.7 Image-to-Video generates videos from images using multi-modal input (text, images, audio, video). Supports first frame generation, first+last frame, and video continuation with 720P/1080P resolution --- ### alibaba/wan-2-7-t2v **URL:** https://inference.sh/apps/alibaba/wan-2-7-t2v Wan 2.7 Text-to-Video generates high-quality videos from text prompts using Alibaba's latest video generation model via DashScope API, supporting 720P/1080P resolution and up to 15 seconds duration --- ### infsh/image-resize **URL:** https://inference.sh/apps/infsh/image-resize Resize images by width, height, scale factor, or megapixel target --- ### pruna/p-image-upscale **URL:** https://inference.sh/apps/pruna/p-image-upscale AI-powered image upscaling up to 128 megapixels with detail and realism enhancement --- ### alibaba/wan-2-7-image-pro **URL:** https://inference.sh/apps/alibaba/wan-2-7-image-pro Wan 2.7 Image Pro is Alibaba's professional image generation model supporting text-to-image, image editing, and multi-reference generation with up to 4K high-definition output --- ### alibaba/wan-2-7-image **URL:** https://inference.sh/apps/alibaba/wan-2-7-image Wan 2.7 Image is Alibaba's fast image generation model supporting text-to-image, image editing, and multi-reference image generation with up to 2K resolution --- # Additional Resources - website: https://inference.sh - docs: https://inference.sh/docs - blog: https://inference.sh/blog - apps: https://inference.sh/apps - github: https://github.com/inference-sh - belt cli: `curl -fsSL https://cli.inference.sh | sh` - python sdk: pip install inferencesh - js sdk: npm install @inferencesh/sdk