IBM: Granite 4.0 Micro
IBM Research · open · 128K context · ₹1/M tokens · Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tune
Every frontier and open-weight model that matters — with context window, INR pricing, language support, and India availability.
IBM Research · open · 128K context · ₹1/M tokens · Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tune
Mistral AI · open · 128K context · ₹2/M tokens · A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, Ger
Inclusionai · open · 256K context · ₹2/M tokens · *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *tok
OpenAI · closed · 391K context · ₹2/M tokens · GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While
Meta · open · 59K context · ₹2/M tokens · Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual
Inference Net · open · 125K context · ₹2/M tokens · Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extracti
OpenAI · closed · 128K context · ₹2/M tokens · gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B act
Alibaba (Qwen) · open-api · 977K context · ₹2/M tokens · Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with stren
Amazon · closed · 125K context · ₹3/M tokens · Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context len
DeepSeek · open · 1024K context · ₹3/M tokens · DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-tok
Cohere · closed · 125K context · ₹3/M tokens · Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiri
~Deepseek · open-api · 1280K context · ₹3/M tokens · This model always redirects to the latest model in the DeepSeek V4 Flash family.
Inception · open-api · 254K context · ₹3/M tokens · Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces an
Sao10K · open · 8K context · ₹3/M tokens · Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balance creativity with impr
DeepSeek · open · 1280K context · ₹3/M tokens · DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited
Tencent · open · 8K context · ₹4/M tokens · Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with
Alibaba (Qwen) · open · 256K context · ₹4/M tokens · Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thin
Google · closed · 1024K context · ₹4/M tokens · Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved through
Meta · open · 128K context · ₹4/M tokens · Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reas
OpenAI · closed · 391K context · ₹4/M tokens · GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While
Meta · open · 128K context · ₹4/M tokens · Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated s
Google · closed · 128K context · ₹4/M tokens · Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 language
Google · closed · 128K context · ₹4/M tokens · Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 language
Mistral AI · open · 32K context · ₹4/M tokens · Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it featur
Inference Net · open · 125K context · ₹4/M tokens · Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Ex
OpenAI · closed · 1023K context · ₹4/M tokens · For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size wit
NVIDIA · open · 256K context · ₹5/M tokens · NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems
IBM Research · open · 128K context · ₹5/M tokens · Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-
Poolside · open · 256K context · ₹5/M tokens · Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (re
Amazon · closed · 293K context · ₹5/M tokens · Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. A
Inclusionai · open · 128K context · ₹5/M tokens · Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native
Inclusionai · open-api · 256K context · ₹5/M tokens · Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is
Z Ai · open · 195K context · ₹5/M tokens · As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, str
Alibaba (Qwen) · open-api · 977K context · ₹5/M tokens · The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts
Alibaba (Qwen) · open · 256K context · ₹6/M tokens · Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code genera
Microsoft · closed · 16K context · ₹6/M tokens · [Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or w
NVIDIA · open · 256K context · ₹6/M tokens · NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agen
Tencent · open · 8K context · ₹6/M tokens · Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows f
Tencent · open · 8K context · ₹6/M tokens · Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs,
~Z Ai · open-api · 1280K context · ₹6/M tokens · This model always redirects to the latest model in the GLM Flash family.
Prism Ml · open · 256K context · ₹6/M tokens · Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding wi
Z Ai · open · 1024K context · ₹6/M tokens · GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention arc
OpenAI · closed · 128K context · ₹6/M tokens · gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lowe
Bytedance Seed · open-api · 256K context · ₹6/M tokens · Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context wind
Mistral AI · open · 256K context · ₹6/M tokens · A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
OpenAI · closed · 125K context · ₹6/M tokens · GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced
Mistral AI · open · 256K context · ₹6/M tokens · Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It com
Alibaba (Qwen) · open · 128K context · ₹7/M tokens · Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seaml
Gryphe · open · 8K context · ₹7/M tokens · One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge
NVIDIA · open · 256K context · ₹7/M tokens · NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-a
Google · closed · 128K context · ₹7/M tokens · Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 language
Tencent · open · 256K context · ₹7/M tokens · Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-worl
Alibaba (Qwen) · open · 256K context · ₹7/M tokens · Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active param
Poolside · open · 1024K context · ₹8/M tokens · Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, s
Alibaba (Qwen) · open · 256K context · ₹8/M tokens · Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targe
Google · closed · 256K context · ₹8/M tokens · Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token dur
Google · closed · 256K context · ₹8/M tokens · Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, c
Upstage · open-api · 512K context · ₹8/M tokens · Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with st
Z Ai · open · 1280K context · ₹8/M tokens · GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention arc
Mistral AI · open · 250K context · ₹8/M tokens · Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved funct
Meta · open · 128K context · ₹8/M tokens · The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instr
Rekaai · open · 16K context · ₹8/M tokens · Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized
Mistral AI · open · 32K context · ₹8/M tokens · Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It e
OpenAI · closed · 1025K context · ₹8/M tokens · GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher
Google · closed · 1024K context · ₹8/M tokens · Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved through
Alibaba (Qwen) · open · 256K context · ₹8/M tokens · Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-p
Meta · open-api · 1024K context · ₹8/M tokens · Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, m
Bytedance · open · 125K context · ₹8/M tokens · UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. B
OpenAI · closed · 1025K context · ₹8/M tokens · GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and
Stepfun · open · 256K context · ₹8/M tokens · Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11
Alibaba (Qwen) · open · 32K context · ₹8/M tokens · Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has grea
Rekaai · open · 64K context · ₹8/M tokens · Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks
Meta · open-api · 1024K context · ₹8/M tokens · Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheap
OpenAI · closed · 391K context · ₹8/M tokens · GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and
Meta · open · 1280K context · ₹8/M tokens · Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It sup
Bytedance Seed · open-api · 256K context · ₹8/M tokens · Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It deliver
Mistral AI · open · 128K context · ₹8/M tokens · The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
OpenAI · closed · 1023K context · ₹8/M tokens · For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size wit
Alibaba (Qwen) · open · 128K context · ₹9/M tokens · Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video.
DeepSeek · open · 1024K context · ₹9/M tokens · DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited
DeepSeek · open · 1024K context · ₹9/M tokens · DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from D
Alibaba (Qwen) · open · 256K context · ₹10/M tokens · Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, an
Alibaba (Qwen) · open · 128K context · ₹10/M tokens · Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seam
Alibaba (Qwen) · open · 128K context · ₹10/M tokens · Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamle
Alibaba (Qwen) · open · 128K context · ₹10/M tokens · Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, mult
Alibaba (Qwen) · open · 256K context · ₹10/M tokens · Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total
OpenAI · closed · 391K context · ₹10/M tokens · GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefi
Google · closed · 1024K context · ₹10/M tokens · Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, a
Alibaba (Qwen) · open · 256K context · ₹11/M tokens · Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimi
Z Ai · open · 128K context · ₹11/M tokens · GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixtu
~Deepseek · open-api · 1024K context · ₹11/M tokens · This model always redirects to the latest model in the DeepSeek Flash family.
Xiaomi · open · 1025K context · ₹12/M tokens · MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in
Tencent · open · 128K context · ₹12/M tokens · Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoni
Mistral AI · open-api · 250K context · ₹13/M tokens · Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middl
Alibaba (Qwen) · open · 977K context · ₹13/M tokens · Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase a
Perceptron · open-api · 32K context · ₹13/M tokens · Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired wit
Alibaba (Qwen) · open · 256K context · ₹13/M tokens · Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard
OpenAI · closed · 128K context · ₹13/M tokens · gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose pro
Mistral AI · open · 256K context · ₹13/M tokens · A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Cohere · closed · 125K context · ₹13/M tokens · command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and
Alibaba (Qwen) · open · 256K context · ₹13/M tokens · Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybr
OpenAI · closed · 128K context · ₹13/M tokens · gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose pro
Mistral AI · open · 256K context · ₹13/M tokens · Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It com
DeepSeek · open · 1024K context · ₹13/M tokens · DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activ
Google · closed · 1024K context · ₹13/M tokens · Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It inclu
OpenAI · closed · 125K context · ₹13/M tokens · GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced
Google · closed · 1024K context · ₹13/M tokens · Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within co
Upstage · open-api · 128K context · ₹13/M tokens · Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers ex
OpenAI · closed · 125K context · ₹13/M tokens · GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced
Alibaba (Qwen) · open · 256K context · ₹14/M tokens · Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-p
Meta · open · 128K context · ₹15/M tokens · Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on con
Tencent · open · 256K context · ₹15/M tokens · Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning lev
Meta · open · 160K context · ₹15/M tokens · Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used
Alibaba (Qwen) · open · 128K context · ₹15/M tokens · Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex s
Meta · open · 1024K context · ₹16/M tokens · Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts a
Alibaba (Qwen) · open-api · 977K context · ₹16/M tokens · Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiere
Alibaba (Qwen) · open · 256K context · ₹16/M tokens · The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and
Alibaba (Qwen) · open-api · 977K context · ₹16/M tokens · Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autono
OpenAI · closed · 391K context · ₹17/M tokens · GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and
Alibaba (Qwen) · open · 80K context · ₹17/M tokens · Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model
Mistral AI · open-api · 32K context · ₹17/M tokens · Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses
Mistral AI · open · 256K context · ₹17/M tokens · The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counter
OpenAI · closed · 1025K context · ₹17/M tokens · GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher
~Openai · open-api · 1025K context · ₹17/M tokens · This model always redirects to the latest model in the GPT Luna family.
Cognitivecomputations · open · 125K context · ₹17/M tokens · Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Veni
OpenAI · closed · 1023K context · ₹17/M tokens · GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context w
NVIDIA · open · 128K context · ₹17/M tokens · NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs
Stepfun · open · 256K context · ₹17/M tokens · Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for na
OpenAI · closed · 1025K context · ₹17/M tokens · GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and
Alibaba (Qwen) · open · 256K context · ₹17/M tokens · Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhanc
Mistral AI · open-api · 128K context · ₹17/M tokens · Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level cap
MiniMax · open · 977K context · ₹17/M tokens · MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion paramet
Alibaba (Qwen) · open · 977K context · ₹17/M tokens · Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long
Alibaba (Qwen) · open · 256K context · ₹18/M tokens · Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instru
DeepSeek · open · 1024K context · ₹18/M tokens · DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from D
Alibaba (Qwen) · open · 128K context · ₹19/M tokens · Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B
Google · closed · 1024K context · ₹21/M tokens · Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro l
Google · closed · 1024K context · ₹21/M tokens · Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, a
Anthropic · closed · 195K context · ₹21/M tokens · Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcem
OpenAI · closed · 391K context · ₹21/M tokens · GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefi
Inception · open-api · 125K context · ₹21/M tokens · Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and re
OpenAI · closed · 16K context · ₹21/M tokens · GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. T
Arcee Ai · open · 256K context · ₹21/M tokens · Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and re
Mistral AI · open-api · 256K context · ₹21/M tokens · Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and re
DeepSeek · open · 160K context · ₹21/M tokens · DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extend
Google · closed · 1024K context · ₹21/M tokens · Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and
DeepSeek · open · 160K context · ₹21/M tokens · DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [Deep
Google · closed · 64K context · ₹21/M tokens · Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and ra
Bytedance Seed · open-api · 256K context · ₹21/M tokens · Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency
Bytedance Seed · open-api · 256K context · ₹21/M tokens · Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context
OpenAI · closed · 391K context · ₹21/M tokens · GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
MiniMax · open · 200K context · ₹21/M tokens · MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 b
Alibaba (Qwen) · open-api · 977K context · ₹22/M tokens · Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.
Alibaba (Qwen) · open-api · 977K context · ₹22/M tokens · The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-expe
Alibaba (Qwen) · open · 256K context · ₹22/M tokens · The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-exper
Alibaba (Qwen) · open-api · 977K context · ₹22/M tokens · Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.
DeepSeek · open · 160K context · ₹22/M tokens · DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduce
DeepSeek · open · 160K context · ₹23/M tokens · DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues re
MiniMax · open · 200K context · ₹23/M tokens · MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments,
DeepSeek · open · 160K context · ₹23/M tokens · DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces Deep
Z Ai · open · 128K context · ₹25/M tokens · GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It su
Meituan · open · 1024K context · ₹25/M tokens · LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level
Google · closed · 1024K context · ₹25/M tokens · Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It inclu
Alibaba (Qwen) · open · 256K context · ₹25/M tokens · Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as
Google · closed · 1024K context · ₹25/M tokens · Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within co
Alibaba (Qwen) · open · 256K context · ₹25/M tokens · Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — a
Meta · open · 128K context · ₹25/M tokens · Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on con
Amazon · closed · 977K context · ₹25/M tokens · Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrate
MiniMax · open-api · 64K context · ₹25/M tokens · MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations. Designed t
Thedrummer · open · 128K context · ₹25/M tokens · Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.
MiniMax · open · 512K context · ₹25/M tokens · MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited f
Google · closed · 32K context · ₹25/M tokens · Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is c
MiniMax · open · 200K context · ₹25/M tokens · MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 b
Alibaba (Qwen) · open-api · 977K context · ₹25/M tokens · Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M t
Mistral AI · open-api · 250K context · ₹25/M tokens · Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middl
Kwaipilot · open-api · 256K context · ₹25/M tokens · KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integra
MiniMax · open · 1024K context · ₹25/M tokens · MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited f
MiniMax · open · 200K context · ₹25/M tokens · MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participat
Alibaba (Qwen) · open · 256K context · ₹26/M tokens · The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixtur
Alibaba (Qwen) · open-api · 977K context · ₹27/M tokens · Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities
DeepSeek · open · 160K context · ₹27/M tokens · DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on ne
Alibaba (Qwen) · open-api · 977K context · ₹27/M tokens · Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and h
Undi95 · open · 6K context · ₹29/M tokens · A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge
Mistral AI · open · 125K context · ₹29/M tokens · Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provi
Alibaba (Qwen) · open · 32K context · ₹30/M tokens · Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has gre
Z Ai · open-api · 1024K context · ₹31/M tokens · GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same
OpenAI · closed · 391K context · ₹31/M tokens · GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image input
Google · closed · 1024K context · ₹31/M tokens · Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require r
Google · closed · 1024K context · ₹31/M tokens · Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs w
Google · closed · 1024K context · ₹31/M tokens · Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reaso
Mistral AI · open-api · 128K context · ₹33/M tokens · Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost
Mancer · open-api · 8K context · ₹33/M tokens · An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.
Alibaba (Qwen) · open · 128K context · ₹33/M tokens · Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is o
Mistral AI · open · 256K context · ₹33/M tokens · Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256
MiniMax · open-api · 977K context · ₹33/M tokens · MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (
Z Ai · open · 200K context · ₹33/M tokens · GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.
Mistral AI · open-api · 128K context · ₹33/M tokens · Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level cap
Thedrummer · open · 1000K context · ₹33/M tokens · UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.
Meta · open · 128K context · ₹33/M tokens · Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usec
OpenAI · closed · 1023K context · ₹33/M tokens · GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context w
Krutrim (Ola) · open-api · 128K context · ₹35/M tokens · Ola Krutrim's 22-language Indic foundation model. INR-priced API with developer-tier free credits. Wide Indic coverage.
Baidu · open · 120K context · ₹35/M tokens · ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token.
DeepSeek · open · 1024K context · ₹35/M tokens · DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context w
Z Ai · open · 200K context · ₹36/M tokens · Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, en
Xiaomi · open · 1025K context · ₹36/M tokens · MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, w
Thinkingmachines · open · 1024K context · ₹38/M tokens · Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned a
Moonshot AI · open · 256K context · ₹38/M tokens · Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi
Alibaba (Qwen) · open · 128K context · ₹38/M tokens · Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching
Sarvam AI · open-api · 32K context · ₹40/M tokens · Indic-native foundation model from Sarvam AI (Bengaluru). INR billing, GST included, Mumbai latency under 100ms. 11 Indian languages with native-quality output.
DeepSeek · open · 160K context · ₹42/M tokens · May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reaso
Bytedance Seed · open-api · 256K context · ₹42/M tokens · Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step
Anthropic · closed · 195K context · ₹42/M tokens · Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude m
Google · closed · 1024K context · ₹42/M tokens · Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro l
Google · closed · 128K context · ₹42/M tokens · Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at F
Google · closed · 64K context · ₹42/M tokens · Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual qual
OpenAI · closed · 16K context · ₹42/M tokens · GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. T
Bytedance Seed · open-api · 256K context · ₹42/M tokens · Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-age
DeepSeek · open · 1024K context · ₹44/M tokens · DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
~Deepseek · open-api · 1024K context · ₹44/M tokens · This model always redirects to the latest model in the DeepSeek Pro family.
Alibaba (Qwen) · open · 256K context · ₹46/M tokens · The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-o
Thedrummer · open · 32K context · ₹46/M tokens · Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent stor
OpenAI · closed · 195K context · ₹46/M tokens · OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model sup
OpenAI · closed · 195K context · ₹46/M tokens · OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabi
Moonshot AI · open · 128K context · ₹48/M tokens · Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active
Writer · open-api · 1016K context · ₹50/M tokens · Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and effic
Z Ai · open · 200K context · ₹50/M tokens · GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it de
NVIDIA · open · 256K context · ₹50/M tokens · NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid
OpenAI · closed · 125K context · ₹50/M tokens · A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. In
Moonshot AI · open · 256K context · ₹50/M tokens · Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI
Z Ai · open · 128K context · ₹50/M tokens · GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a
Moonshot AI · open · 256K context · ₹50/M tokens · Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillio
Z Ai · open · 64K context · ₹50/M tokens · GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B
Microsoft · closed · 64K context · ₹52/M tokens · WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it con
OpenAI · closed · 391K context · ₹52/M tokens · GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that requi
OpenAI · closed · 391K context · ₹52/M tokens · GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural
Google · closed · 1024K context · ₹52/M tokens · Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilit
Z Ai · open · 1024K context · ₹54/M tokens · GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workf
Google · closed · 8K context · ₹54/M tokens · Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-s
Alibaba (Qwen) · open-api · 977K context · ₹54/M tokens · Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous progr
Sao10K · open · 128K context · ₹54/M tokens · Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3
DeepSeek · open · 1024K context · ₹55/M tokens · DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
Alibaba (Qwen) · open · 32K context · ₹55/M tokens · Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upo
Z Ai · open · 1024K context · ₹58/M tokens · GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with
DeepSeek · open · 62K context · ₹58/M tokens · DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with
Z Ai · open · 1024K context · ₹58/M tokens · GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workf
Nous Research · open · 128K context · ₹58/M tokens · Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic ca
Aion Labs · open-api · 128K context · ₹58/M tokens · Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation pro
Moonshot AI · open · 256K context · ₹59/M tokens · MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts
Kwaipilot · open-api · 256K context · ₹62/M tokens · KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to aut
Google · closed · 1024K context · ₹63/M tokens · Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs w
Google · closed · 1024K context · ₹63/M tokens · Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require r
Mistral AI · open-api · 256K context · ₹63/M tokens · Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic
OpenAI · closed · 391K context · ₹63/M tokens · GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image input
Google · closed · 1024K context · ₹63/M tokens · Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized
~Openai · open-api · 391K context · ₹63/M tokens · This model always redirects to the latest model in the GPT Mini family.
Google · closed · 1024K context · ₹63/M tokens · Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reaso
~Google · open-api · 1024K context · ₹63/M tokens · This model always redirects to the latest model in the Gemini Flash family.
Alibaba (Qwen) · open-api · 256K context · ₹65/M tokens · Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail
Alibaba (Qwen) · open-api · 256K context · ₹65/M tokens · Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By sig
DeepSeek · open · 8K context · ₹67/M tokens · DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [Dee
Morph · open-api · 80K context · ₹67/M tokens · Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the foll
Alibaba (Qwen) · open · 125K context · ₹67/M tokens · Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, g
Aion Labs · open-api · 128K context · ₹67/M tokens · Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and confl
Amazon · closed · 293K context · ₹67/M tokens · Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of D
Aion Labs · open-api · 32K context · ₹67/M tokens · Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, whe
~Z Ai · open-api · 1280K context · ₹69/M tokens · This model always redirects to the latest GLM model from Z.ai.
Tencent · open · 1024K context · ₹70/M tokens · Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-us
Sao10K · open · 128K context · ₹71/M tokens · Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao1
Relace · open-api · 250K context · ₹71/M tokens · Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and
OpenAI · closed · 391K context · ₹73/M tokens · GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reas
Morph · open-api · 256K context · ₹75/M tokens · Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to
Z Ai · open · 1280K context · ₹76/M tokens · GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with
Sakana · open-api · 256K context · ₹79/M tokens · Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts.
Moonshot AI · open · 256K context · ₹79/M tokens · Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It
Z Ai · open · 200K context · ₹81/M tokens · GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minu
Thinkingmachines · open · 512K context · ₹84/M tokens · Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for gener
OpenAI · closed · 1025K context · ₹84/M tokens · GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for hig
xAI · closed · 977K context · ₹84/M tokens · Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks
Anthropic · closed · 977K context · ₹84/M tokens · Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking w
Relace · open-api · 250K context · ₹84/M tokens · The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to R
OpenAI · closed · 1025K context · ₹84/M tokens · GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyd
Nous Research · open · 128K context · ₹84/M tokens · Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi
OpenAI · closed · 1025K context · ₹84/M tokens · GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-qu
OpenAI · closed · 4K context · ₹84/M tokens · GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. T
xAI · closed · 250K context · ₹84/M tokens · Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text outp
~Anthropic · open-api · 195K context · ₹84/M tokens · This model always redirects to the latest model in the Claude Haiku family.
Google · closed · 1024K context · ₹84/M tokens · Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more effici
OpenAI · closed · 1025K context · ₹84/M tokens · GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at c
Thinkingmachines · open · 1024K context · ₹84/M tokens · Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for gener
Perplexity · open-api · 124K context · ₹84/M tokens · Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for companies seeking t
OpenAI · closed · 1023K context · ₹84/M tokens · GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It support
Nous Research · open · 128K context · ₹84/M tokens · Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can
Anthropic · closed · 195K context · ₹84/M tokens · Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude m
OpenAI · closed · 195K context · ₹84/M tokens · o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technica
Alibaba (Qwen) · open-api · 256K context · ₹86/M tokens · Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total par
OpenAI · closed · 195K context · ₹92/M tokens · OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-seri
OpenAI · closed · 195K context · ₹92/M tokens · OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabi
OpenAI · closed · 195K context · ₹92/M tokens · OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model sup
OpenAI · closed · 195K context · ₹92/M tokens · OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for
Z Ai · open-api · 198K context · ₹100/M tokens · GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, a
Z Ai · open-api · 198K context · ₹100/M tokens · GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply op
xAI · closed · 1953K context · ₹104/M tokens · Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct dee
OpenAI · closed · 125K context · ₹104/M tokens · GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turb
OpenAI · closed · 391K context · ₹104/M tokens · GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version
Meta · open-api · 1024K context · ₹104/M tokens · Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, wi
Google · closed · 1024K context · ₹104/M tokens · Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilit
Meta · open-api · 1024K context · ₹104/M tokens · Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and o
Google · closed · 1024K context · ₹104/M tokens · Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilit
OpenAI · closed · 1025K context · ₹104/M tokens · GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K outpu
Meta · open-api · 1024K context · ₹104/M tokens · Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of informatio
OpenAI · closed · 391K context · ₹104/M tokens · GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessi
OpenAI · closed · 391K context · ₹104/M tokens · GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural
xAI · closed · 1953K context · ₹104/M tokens · Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the
OpenAI · closed · 391K context · ₹104/M tokens · GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that requi
xAI · closed · 977K context · ₹104/M tokens · Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks
Alibaba (Qwen) · open-api · 977K context · ₹123/M tokens · Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular st
Google · closed · 1024K context · ₹125/M tokens · Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized
Anthropic · closed · 977K context · ₹125/M tokens · Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative de
Anthropic · closed · 977K context · ₹125/M tokens · Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performa
Mistral AI · open-api · 256K context · ₹125/M tokens · Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic
OpenAI · closed · 4K context · ₹125/M tokens · This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.
~Moonshotai · open-api · 1024K context · ₹142/M tokens · This model always redirects to the latest model in the Kimi family.
Moonshot AI · open · 1024K context · ₹142/M tokens · Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic w
OpenAI · closed · 391K context · ₹146/M tokens · GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasonin
OpenAI · closed · 391K context · ₹146/M tokens · GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reas
OpenAI · closed · 125K context · ₹146/M tokens · GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It use
OpenAI · closed · 391K context · ₹146/M tokens · GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development s
Anthropic · closed · 977K context · ₹167/M tokens · Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking w
Sakana · open-api · 977K context · ₹167/M tokens · Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a
Alibaba (Qwen) · open · 986K context · ₹167/M tokens · Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion a
OpenAI · closed · 1025K context · ₹167/M tokens · GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-qu
Mistral AI · open-api · 125K context · ₹167/M tokens · This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, J
Perplexity · open-api · 125K context · ₹167/M tokens · Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-re
xAI · closed · 488K context · ₹167/M tokens · Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Google · closed · 64K context · ₹167/M tokens · Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly impr
Google · closed · 128K context · ₹167/M tokens · Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly impr
OpenAI · closed · 1025K context · ₹167/M tokens · GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for hig
OpenAI · closed · 195K context · ₹167/M tokens · o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technica
OpenAI · closed · 1025K context · ₹167/M tokens · GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyd
~Openai · open-api · 1025K context · ₹167/M tokens · This model always redirects to the latest model in the GPT Sol family.
Mistral AI · open-api · 128K context · ₹167/M tokens · This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSO
xAI · closed · 488K context · ₹167/M tokens · Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
~Anthropic · open-api · 977K context · ₹167/M tokens · This model always redirects to the latest model in the Claude Sonnet family.
OpenAI · closed · 1023K context · ₹167/M tokens · GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It support
Google · closed · 1024K context · ₹167/M tokens · Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more
Perplexity · open-api · 125K context · ₹167/M tokens · Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously searches, rea
Alibaba (Qwen) · open-api · 977K context · ₹167/M tokens · Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, imag
Mistral AI · open · 64K context · ₹167/M tokens · Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparall
OpenAI · closed · 1025K context · ₹167/M tokens · GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at c
Google · closed · 1024K context · ₹167/M tokens · Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more effici
~Google · open-api · 1024K context · ₹167/M tokens · This model always redirects to the latest model in the Gemini Pro family.
~Openai · open-api · 1025K context · ₹167/M tokens · This model always redirects to the latest model in the GPT Terra family.
~X Ai · open-api · 488K context · ₹167/M tokens · This model always redirects to the latest Grok model from xAI.
Alibaba (Qwen) · open · 1024K context · ₹167/M tokens · Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion a
Anthropic · closed · 977K context · ₹209/M tokens · Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tas
OpenAI · closed · 125K context · ₹209/M tokens · The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readabili
OpenAI · closed · 391K context · ₹209/M tokens · GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for e
OpenAI · closed · 125K context · ₹209/M tokens · The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [h
OpenAI · closed · 1025K context · ₹209/M tokens · GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K outpu
Anthropic · closed · 195K context · ₹209/M tokens · Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers
OpenAI · closed · 1025K context · ₹209/M tokens · GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved to
Anthropic · closed · 977K context · ₹209/M tokens · Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.
Unbiased · open-api · 256K context · ₹209/M tokens · Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of ge
Cohere · closed · 250K context · ₹209/M tokens · Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding us
OpenAI · closed · 125K context · ₹209/M tokens · GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turb
Cohere · closed · 125K context · ₹209/M tokens · command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to
OpenAI · closed · 125K context · ₹209/M tokens · The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and mainta
Anthropic · closed · 977K context · ₹209/M tokens · Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than
Anthropic · closed · 977K context · ₹209/M tokens · Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reason
Anthracite Org · open · 32K context · ₹209/M tokens · This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet
Amazon · closed · 977K context · ₹209/M tokens · Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.
Aion Labs · open-api · 128K context · ₹250/M tokens · Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in wh
Anthropic · closed · 977K context · ₹250/M tokens · Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and
Perplexity · open-api · 195K context · ₹250/M tokens · Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reas
Moonshot AI · open · 1024K context · ₹250/M tokens · Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic w
Anthropic · closed · 977K context · ₹250/M tokens · Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performa
OpenAI · closed · 16K context · ₹250/M tokens · This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Tr
Perplexity · open-api · 195K context · ₹250/M tokens · Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-re
Anthropic · closed · 977K context · ₹250/M tokens · Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative de
OpenAI · closed · 1025K context · ₹418/M tokens · GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved to
Anthropic · closed · 977K context · ₹418/M tokens · Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reason
Anthropic · closed · 977K context · ₹418/M tokens · Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than
OpenAI · closed · 125K context · ₹418/M tokens · The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
OpenAI · closed · 391K context · ₹418/M tokens · GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new
OpenAI · closed · 1025K context · ₹418/M tokens · GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-qu
OpenAI · closed · 1025K context · ₹418/M tokens · GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work,
Anthropic · closed · 977K context · ₹418/M tokens · Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tas
Sakana · open-api · 977K context · ₹418/M tokens · Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system
Anthropic · closed · 977K context · ₹418/M tokens · Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.
Anthropic · closed · 977K context · ₹418/M tokens · Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long
Anthropic · closed · 977K context · ₹418/M tokens · Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output
~Anthropic · open-api · 977K context · ₹418/M tokens · This model always redirects to the latest model in the Claude Opus family.
OpenAI · closed · 125K context · ₹418/M tokens · GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turb
Sakana · open-api · 977K context · ₹418/M tokens · Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration sys
Anthropic · closed · 195K context · ₹418/M tokens · Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers
OpenAI · closed · 391K context · ₹626/M tokens · GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that r
Anthropic · closed · 195K context · ₹626/M tokens · Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on
OpenAI · closed · 266K context · ₹668/M tokens · [GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It
OpenAI · closed · 1025K context · ₹835/M tokens · GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work,
OpenAI · closed · 391K context · ₹835/M tokens · [GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvement
~Openai · open-api · 1025K context · ₹835/M tokens · This model always redirects to the latest model in the GPT Astra family.
~Anthropic · open-api · 977K context · ₹835/M tokens · This model always redirects to the latest model in the Claude Fable family.
Anthropic · closed · 977K context · ₹835/M tokens · Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output
Anthropic · closed · 977K context · ₹835/M tokens · Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long
OpenAI · closed · 125K context · ₹835/M tokens · The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
OpenAI · closed · 1025K context · ₹835/M tokens · GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-qu
OpenAI · closed · 391K context · ₹877/M tokens · GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for comp
OpenAI · closed · 195K context · ₹1,252/M tokens · The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale
Anthropic · closed · 195K context · ₹1,252/M tokens · Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workf
OpenAI · closed · 1025K context · ₹1,252/M tokens · GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It
OpenAI · closed · 391K context · ₹1,252/M tokens · GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that r
Anthropic · closed · 195K context · ₹1,252/M tokens · Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on
OpenAI · closed · 1025K context · ₹1,252/M tokens · GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context windo
OpenAI · closed · 195K context · ₹1,670/M tokens · The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to
OpenAI · closed · 391K context · ₹1,754/M tokens · GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for comp
OpenAI · closed · 1025K context · ₹2,505/M tokens · GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context windo
OpenAI · closed · 1025K context · ₹2,505/M tokens · GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It
OpenAI · closed · 8K context · ₹2,505/M tokens · OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due t
OpenAI · closed · 195K context · ₹12,525/M tokens · The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to
OpenRouter · open-api · 1953K context · The experimental version of our Auto Router where we test new improvements. Use it to get the latest and greatest version of our general purpose auto router, bu
Google · closed · 1024K context · Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can g
OpenRouter · open-api · 1953K context · The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set
OpenRouter · open-api · 977K context · Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fe
OpenRouter · open-api · 1953K context · The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community
Google · closed · 1024K context · 30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, yo
OpenRouter · open-api · 125K context · Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder w
OpenRouter · open-api · 195K context · The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smart
MeitY · Bhashini · open-api · 4K context · Government of India's national-mission translation service covering 22 Indian languages. Free tier for non-commercial; commercial usage via Bhashadaan partner n