Skip to content
  • Models
  • Rankings
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for openai
    OpenAI: GPT-6 Astra (batch)GPT-6 Astra (batch)Batch variant
    42.9M tokens

    GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

    by openaiSep 4, 20261.05M context$5/M input tokens$25/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 AstraGPT-6 Astra
    59.5B tokens

    GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

    by openaiSep 4, 20261.05M context$10/M input tokens$50/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Astra Pro (batch)GPT-6 Astra Pro (batch)Batch variant
    1.76M tokens

    GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 4, 20261.05M context$5/M input tokens$25/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Astra ProGPT-6 Astra Pro
    13.8B tokens

    GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 4, 20261.05M context$10/M input tokens$50/M output tokens
  • Favicon for microsoft
    Microsoft: MAI-Image-2.6MAI-Image-2.6
    11M tokens

    Microsoft's MAI-Image-2.6 is an image generation and editing model available via Azure AI Foundry. It creates images from text prompts and supports image-guided editing across multiple aspect ratios.

    by microsoftSep 4, 20264K context$5/M tokens
  • Favicon for microsoft
    Microsoft: MAI-Image-2.6 FlashMAI-Image-2.6 Flash
    9.21M tokens

    Microsoft's MAI-Image-2.6 Flash is the lower-latency variant of MAI-Image-2.6, available via Azure AI Foundry. It supports image generation and image-guided editing across multiple aspect ratios.

    by microsoftSep 4, 20264K context$1.75/M tokens
  • Favicon for inclusionai
    inclusionAI: Ling 3.0 Flash Sante (free)Ling 3.0 Flash Sante (free)Free variant
    28.6B tokens

    Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for medical knowledge reasoning, clinical safety, evidence-based retrieval, and long-horizon medical tasks, while retaining general capabilities in reasoning, coding, and agentic tasks.

    by inclusionaiSep 4, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 Max (0902)Qwen3.8 Max (0902)
    30.7B tokens

    Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. This snapshot is post-trained for coding and agentic work, including multi-step software projects, multi-tool orchestration, and long-horizon task execution. It also targets chart reasoning, document parsing, and multimodal understanding over long documents and extended video. Tool calling, structured outputs, and configurable reasoning effort are supported.

    by qwenSep 3, 20261M context$2/M input tokens$6/M output tokens
  • Favicon for microsoft
    Microsoft: MAI-Transcribe 2MAI-Transcribe 2
    112M characters

    MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech, speaker diarization, word-level timestamps, keyword biasing for domain-specific terminology, and configurable verbatim or clean transcription styles. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, and is faster than MAI-Transcribe-1.5 on long-form audio. On OpenRouter, set response_format to "verbose_json" for segment timestamps, and add timestamp_granularities: ["word"] for word-level timestamps. Set provider.options.azure.diarization.enabled to true for speaker labels, provide keywords through provider.options.azure.phraseList.phrases, and select a transcription style through provider.options.azure.enhancedMode.modelOptions.transcribeStyle. See the speech-to-text guide.

    by microsoftSep 3, 2026$0.10/hour
  • Favicon for meta
    Meta: Muse Spark 1.3 ContributorMuse Spark 1.3 Contributor
    18+
    530B tokens

    Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed. Prompts and outputs may be used to improve Meta’s products.

    by metaSep 2, 20261.05M context$0.10/M input tokens$0.20/M output tokens
  • Favicon for meta
    Meta: Muse Spark 1.3Muse Spark 1.3
    18+
    110B tokens

    Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed, with an emphasis on concise execution.

    by metaSep 2, 20261.05M context$1.25/M input tokens$4.25/M output tokens
  • Favicon for google
    Google: Gemini 3.8 FlashGemini 3.8 Flash
    50% off
    830B tokens

    Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

    by googleSep 2, 20261.05M context$0.75/M input tokens$3.75/M output tokens
  • Favicon for google
    Google: Gemini 3.8 Flash (batch)Gemini 3.8 Flash (batch)
    50% off
    Batch variant
    1.14B tokens

    Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

    by googleSep 2, 20261.05M context$0.375/M input tokens$1.875/M output tokens
  • Favicon for minimax
    MiniMax: H3 MaxH3 Max
    10 hours

    MiniMax H3 Max is a video-generation model from MiniMax, jointly released with fal.ai. Derived through additional training from MiniMax H3, it is designed for faster text-to-video and image-to-video generation with controlled first-frame or last-frame keyframes.

    by minimaxSep 2, 2026from $0.05/second
  • Favicon for anthropic
    Anthropic: Claude Fable 5.1 (batch)Claude Fable 5.1 (batch)Batch variant
    34M tokens

    Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks.

    by anthropicSep 1, 20261M context$5/M input tokens$25/M output tokens
  • Favicon for anthropic
    Anthropic: Claude Fable 5.1Claude Fable 5.1
    252B tokens

    Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks.

    by anthropicSep 1, 20261M context$10/M input tokens$50/M output tokens
  • Favicon for inception
    Inception: Mercury 2.5 PreviewMercury 2.5 Preview
    80% off
    38.5B tokens

    Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.

    by inceptionAug 31, 2026260K context$0.04/M input tokens$0.15/M output tokens
  • Favicon for ibm-granite
    IBM: Granite 4.2 8BGranite 4.2 8B
    2.04B tokens

    Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort, and non-thinking modes. The model supports 12 languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese.

    by ibm-graniteAug 31, 2026131K context$0.06/M input tokens$0.25/M output tokens
  • Favicon for tencent
    Tencent: Hy4 previewHy4 preview
    15.5T tokens

    Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

    by tencentAug 28, 20261.05M context$0.834/M input tokens$2.501/M output tokens
  • Favicon for alibaba
    Alibaba: Wan 3.0 PrimeWan 3.0 Prime
    4 hours

    Wan 3.0 Prime is a fast-mode variant of Wan 3.0 from Alibaba. It supports text-to-video and first-frame image-to-video generation.

    by alibabaAug 27, 2026from $0.068/second