AI MODEL DIRECTORY

Model Wiki

Explore major large language models with capability notes, technical parameters, use cases and trade-offs

0models0Provider0Family

domainOpenAI60 models

OpenAIgpt

GPT-5.6 Sol

OpenAI flagship GPT-5.6 model for complex reasoning, coding, biology and cybersecurity workflows

GPT-5.6 Sol is an OpenAI general-purpose model for demanding reasoning, coding agents, biology analysis, security research and highest-capability workflows.

1050K contextMultimodalInput $5/1M tokensOutput $30/1M tokensReleased June 2026
OpenAIgpt

GPT-5.6 Terra

OpenAI balanced GPT-5.6 model for everyday work, coding and high-quality general tasks

GPT-5.6 Terra is an OpenAI general-purpose model for everyday assistants, coding collaboration, content generation, complex Q&A and balanced production workloads.

1050K contextMultimodalInput $2.5/1M tokensOutput $15/1M tokensReleased June 2026
OpenAIgpt-mini

GPT-5.6 Luna

OpenAI lightweight GPT-5.6 model for low-cost, high-volume and fast-response workloads

GPT-5.6 Luna is an OpenAI general-purpose model for high-volume support, batch processing, lightweight automation, summarization and cost-sensitive fast tasks.

1050K contextMultimodalInput $1/1M tokensOutput $6/1M tokensReleased June 2026
OpenAIgpt-pro

GPT-5.5 Pro

GPT-5.5 Pro: a gpt-pro model from OpenAI, ~1.1M context, knowledge cutoff 2025-12-01

GPT-5.5 Pro is a gpt-pro model from OpenAI (~1.1M context, input around $30/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1050K contextMultimodalInput $30/1M tokensOutput $180/1M tokensReleased April 2026
OpenAIgpt

GPT-5.5

OpenAI flagship general-purpose model for complex reasoning, coding and high-quality generation

GPT-5.5 is an OpenAI general-purpose model for hard Q&A, complex coding, long-form analysis and agent workflows.

1050K contextMultimodalInput $5/1M tokensOutput $30/1M tokensReleased April 2026
OpenAIgpt-pro

GPT-5.4 Pro

GPT-5.4 Pro: a gpt-pro model from OpenAI, ~1.1M context, knowledge cutoff 2025-08-31

GPT-5.4 Pro is a gpt-pro model from OpenAI (~1.1M context, input around $30/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1050K contextMultimodalInput $30/1M tokensOutput $180/1M tokensReleased March 2026
OpenAIgpt

GPT-5.3 Chat (latest)

GPT-5.3 Chat (latest): a gpt model from OpenAI, ~128K context, knowledge cutoff 2025-08-31

GPT-5.3 Chat (latest) is a gpt model from OpenAI (~128K context, input around $1.75/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $1.75/1M tokensOutput $14/1M tokensReleased March 2026
OpenAIgpt-codex-spark

GPT-5.3 Codex Spark

GPT-5.3 Codex Spark: a gpt-codex-spark model from OpenAI, ~128K context, knowledge cutoff 2025-08-31

GPT-5.3 Codex Spark is a gpt-codex-spark model from OpenAI (~128K context, input around $1.75/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $1.75/1M tokensOutput $14/1M tokensReleased February 2026
OpenAIgpt-codex

GPT-5.3 Codex

GPT-5.3 Codex: a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2025-08-31

GPT-5.3 Codex is a gpt-codex model from OpenAI (~400K context, input around $1.75/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $1.75/1M tokensOutput $14/1M tokensReleased February 2026
OpenAIgpt

GPT-5.2

GPT-5.2: a gpt model from OpenAI, ~400K context, knowledge cutoff 2025-08-31

GPT-5.2 is a gpt model from OpenAI (~400K context, input around $1.75/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $1.75/1M tokensOutput $14/1M tokensReleased December 2025
OpenAIgpt-pro

GPT-5.2 Pro

GPT-5.2 Pro: a gpt-pro model from OpenAI, ~400K context, knowledge cutoff 2025-08-31

GPT-5.2 Pro is a gpt-pro model from OpenAI (~400K context, input around $21/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $21/1M tokensOutput $168/1M tokensReleased December 2025
OpenAIgpt-codex

GPT-5.2 Chat

GPT-5.2 Chat: a gpt-codex model from OpenAI, ~128K context, knowledge cutoff 2025-08-31

GPT-5.2 Chat is a gpt-codex model from OpenAI (~128K context, input around $1.75/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $1.75/1M tokensOutput $14/1M tokensReleased December 2025
OpenAIgpt-codex

GPT-5.2 Codex

GPT-5.2 Codex: a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2025-08-31

GPT-5.2 Codex is a gpt-codex model from OpenAI (~400K context, input around $1.75/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released December 2025
OpenAIgpt-codex

GPT-5.1 Codex mini

GPT-5.1 Codex mini: a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5.1 Codex mini is a gpt-codex model from OpenAI (~400K context, input around $0.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released November 2025
OpenAIgpt-codex

GPT-5.1 Chat

GPT-5.1 Chat: a gpt-codex model from OpenAI, ~128K context, knowledge cutoff 2024-09-30

GPT-5.1 Chat is a gpt-codex model from OpenAI (~128K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released November 2025
OpenAIgpt

GPT-5.1

GPT-5.1: a gpt model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5.1 is a gpt model from OpenAI (~400K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $1.25/1M tokensOutput $10/1M tokensReleased November 2025
OpenAIgpt-codex

GPT-5.1 Codex Max

GPT-5.1 Codex Max: a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5.1 Codex Max is a gpt-codex model from OpenAI (~400K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released November 2025
OpenAIgpt-codex

GPT-5.1 Codex

GPT-5.1 Codex: a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5.1 Codex is a gpt-codex model from OpenAI (~400K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released November 2025
OpenAIgpt-pro

GPT-5 Pro

GPT-5 Pro: a gpt-pro model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5 Pro is a gpt-pro model from OpenAI (~400K context, input around $15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $15/1M tokensOutput $120/1M tokensReleased October 2025
OpenAIgpt-codex

GPT-5-Codex

GPT-5-Codex: a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5-Codex is a gpt-codex model from OpenAI (~400K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released September 2025
OpenAIgpt-codex

GPT-5 Chat (latest)

GPT-5 Chat (latest): a gpt-codex model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5 Chat (latest) is a gpt-codex model from OpenAI (~400K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released August 2025
OpenAIgpt-nano

GPT-5 Nano

GPT-5 Nano: a gpt-nano model from OpenAI, ~400K context, knowledge cutoff 2024-05-30

GPT-5 Nano is a gpt-nano model from OpenAI (~400K context, input around $0.05/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $0.05/1M tokensOutput $0.4/1M tokensReleased August 2025
OpenAIgpt

GPT-5

GPT-5: a gpt model from OpenAI, ~400K context, knowledge cutoff 2024-09-30

GPT-5 is a gpt model from OpenAI (~400K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

400K contextMultimodalInput $1.25/1M tokensOutput $10/1M tokensReleased August 2025
OpenAIgpt-mini

GPT-5 Mini

GPT-5 Mini model profile for capabilities, pricing and use cases

GPT-5 Mini is a large language model from OpenAI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

400K contextMultimodalInput $0.25/1M tokensOutput $2/1M tokensReleased August 2025
OpenAIo-pro

o3-pro

o3-pro: a o-pro model from OpenAI, ~200K context, knowledge cutoff 2024-05

o3-pro is a o-pro model from OpenAI (~200K context, input around $20/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $20/1M tokensOutput $80/1M tokensReleased June 2025
OpenAIo-mini

o4-mini

o4-mini: a o-mini model from OpenAI, ~200K context, knowledge cutoff 2024-05

o4-mini is a o-mini model from OpenAI (~200K context, input around $1.1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $1.1/1M tokensOutput $4.4/1M tokensReleased April 2025
OpenAIo

o3

o3: a o model from OpenAI, ~200K context, knowledge cutoff 2024-05

o3 is a o model from OpenAI (~200K context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $2/1M tokensOutput $8/1M tokensReleased April 2025
OpenAIgpt

GPT-4.1

GPT-4.1: a gpt model from OpenAI, ~1M context, knowledge cutoff 2024-04

GPT-4.1 is a gpt model from OpenAI (~1M context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1048K contextMultimodalInput $2/1M tokensOutput $8/1M tokensReleased April 2025
OpenAIgpt-mini

GPT-4.1 mini

GPT-4.1 mini: a gpt-mini model from OpenAI, ~1M context, knowledge cutoff 2024-04

GPT-4.1 mini is a gpt-mini model from OpenAI (~1M context, input around $0.4/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1048K contextMultimodalInput $0.4/1M tokensOutput $1.6/1M tokensReleased April 2025
OpenAIgpt-nano

GPT-4.1 nano

GPT-4.1 nano: a gpt-nano model from OpenAI, ~1M context, knowledge cutoff 2024-04

GPT-4.1 nano is a gpt-nano model from OpenAI (~1M context, input around $0.1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1048K contextMultimodalInput $0.1/1M tokensOutput $0.4/1M tokensReleased April 2025
OpenAIo-pro

o1-pro

o1-pro: a o-pro model from OpenAI, ~200K context, knowledge cutoff 2023-09

o1-pro is a o-pro model from OpenAI (~200K context, input around $150/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $150/1M tokensOutput $600/1M tokensReleased March 2025
OpenAIo-mini

o3-mini

o3-mini model profile for capabilities, pricing and use cases

o3-mini is a large language model from OpenAI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

200K contextInput $1.1/1M tokensOutput $4.4/1M tokensReleased December 2024
OpenAIo

o1

o1 model profile for capabilities, pricing and use cases

o1 is a large language model from OpenAI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

200K contextMultimodalInput $15/1M tokensOutput $60/1M tokensReleased December 2024
OpenAIgpt

GPT-4o (2024-11-20)

GPT-4o (2024-11-20): a gpt model from OpenAI, ~128K context, knowledge cutoff 2023-09

GPT-4o (2024-11-20) is a gpt model from OpenAI (~128K context, input around $2.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $2.5/1M tokensOutput $10/1M tokensReleased November 2024
OpenAIo

o1-preview

o1-preview: a o model from OpenAI, ~128K context, knowledge cutoff 2023-09

o1-preview is a o model from OpenAI (~128K context, input around $15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released September 2024
OpenAIo-mini

o1-mini

o1-mini model profile for capabilities, pricing and use cases

o1-mini is a large language model from OpenAI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

Released September 2024
OpenAIgpt

GPT-4o (2024-08-06)

GPT-4o (2024-08-06): a gpt model from OpenAI, ~128K context, knowledge cutoff 2023-09

GPT-4o (2024-08-06) is a gpt model from OpenAI (~128K context, input around $2.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $2.5/1M tokensOutput $10/1M tokensReleased August 2024
OpenAIgpt-mini

GPT-4o mini

Compact OpenAI multimodal model for high-volume, cost-sensitive workloads

GPT-4o mini is a lightweight OpenAI model designed for lower-cost, high-throughput applications while keeping useful multimodal and tool-assisted capabilities.

128K contextMultimodalInput $0.15/1M tokensOutput $0.6/1M tokensReleased July 2024
OpenAIo-mini

o4-mini-deep-research

o4-mini-deep-research: a o-mini model from OpenAI, ~200K context, knowledge cutoff 2024-05

o4-mini-deep-research is a o-mini model from OpenAI (~200K context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released June 2024
OpenAIo

o3-deep-research

o3-deep-research: a o model from OpenAI, ~200K context, knowledge cutoff 2024-05

o3-deep-research is a o model from OpenAI (~200K context, input around $10/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released June 2024
OpenAIgpt

GPT-4o (2024-05-13)

GPT-4o (2024-05-13): a gpt model from OpenAI, ~128K context, knowledge cutoff 2023-09

GPT-4o (2024-05-13) is a gpt model from OpenAI (~128K context, input around $5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $5/1M tokensOutput $15/1M tokensReleased May 2024
OpenAIgpt

GPT-4o

OpenAI flagship multimodal model for text, vision and real-time interaction

GPT-4o is OpenAI's general-purpose flagship multimodal model. It is suitable for products that need strong text understanding, image analysis, code assistance, tool calling and stable conversation quality at scale.

128K contextMultimodalInput $2.5/1M tokensOutput $10/1M tokensReleased May 2024
OpenAItext-embedding

text-embedding-3-large

text-embedding-3-large: a text-embedding model from OpenAI, ~8.2K context, knowledge cutoff 2024-01

text-embedding-3-large is a text-embedding model from OpenAI (~8.2K context, input around $0.13/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $0.13/1M tokensOutput $0/1M tokensReleased January 2024
OpenAItext-embedding

text-embedding-3-small

text-embedding-3-small: a text-embedding model from OpenAI, ~8.2K context, knowledge cutoff 2024-01

text-embedding-3-small is a text-embedding model from OpenAI (~8.2K context, input around $0.02/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $0.02/1M tokensOutput $0/1M tokensReleased January 2024
OpenAIgpt

GPT-4 Turbo

GPT-4 Turbo: a gpt model from OpenAI, ~128K context, knowledge cutoff 2023-12

GPT-4 Turbo is a gpt model from OpenAI (~128K context, input around $10/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $10/1M tokensOutput $30/1M tokensReleased November 2023
OpenAIgpt

GPT-4

GPT-4: a gpt model from OpenAI, ~8.2K context, knowledge cutoff 2023-11

GPT-4 is a gpt model from OpenAI (~8.2K context, input around $30/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $30/1M tokensOutput $60/1M tokensReleased November 2023
OpenAIgpt

GPT-3.5-turbo

GPT-3.5-turbo: a gpt model from OpenAI, ~16.4K context, knowledge cutoff 2021-09-01

GPT-3.5-turbo is a gpt model from OpenAI (~16.4K context, input around $0.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

16K contextInput $0.5/1M tokensOutput $1.5/1M tokensReleased March 2023
OpenAItext-embedding

text-embedding-ada-002

text-embedding-ada-002: a text-embedding model from OpenAI, ~8.2K context, knowledge cutoff 2022-12

text-embedding-ada-002 is a text-embedding model from OpenAI (~8.2K context, input around $0.1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $0.1/1M tokensOutput $0/1M tokensReleased December 2022
OpenAIgpt-mini

GPT-5.4 nano

Tiny GPT-5.4 variant for very low-cost and high-throughput tasks

GPT-5.4 nano is an OpenAI general-purpose model for batch tagging, simple extraction, routing, preprocessing and extremely cost-sensitive workloads.

400K contextMultimodalInput $0.2/1M tokensOutput $1.25/1M tokens
OpenAIgpt

GPT-5.4

High-capability OpenAI general model balancing quality and cost

GPT-5.4 is an OpenAI general-purpose model for product assistants, content generation, coding assistance, structured processing and multimodal understanding.

1050K contextMultimodalInput $2.5/1M tokensOutput $15/1M tokens
OpenAIgpt-mini

GPT-5.4 mini

Lower-latency and lower-cost GPT-5.4 variant

GPT-5.4 mini is an OpenAI general-purpose model for high-volume support, summarization, rewriting, lightweight classification and cost-sensitive automation.

400K contextMultimodalInput $0.75/1M tokensOutput $4.5/1M tokens
OpenAIgpt-image

GPT Image 2

OpenAI image generation and editing model for high-quality visual creation

GPT Image 2 is an OpenAI image model for text-to-image generation, image editing, creative design and visual content workflows.

MultimodalInput $5/1M tokensOutput $30/1M tokens
OpenAIrealtime

GPT Realtime 2

Reasoning-focused realtime voice model for low-latency audio interactions

GPT Realtime 2 is an OpenAI realtime model for low-latency voice input, voice output and interactive conversational experiences.

OpenAIrealtime-mini

GPT Realtime mini

Cost-efficient OpenAI realtime model for voice applications

GPT Realtime mini is an OpenAI realtime model for low-latency voice input, voice output and interactive conversational experiences.

OpenAIrealtime

GPT Realtime 1.5

OpenAI realtime voice model for audio input and audio output

GPT Realtime 1.5 is an OpenAI realtime model for low-latency voice input, voice output and interactive conversational experiences.

OpenAIrealtime

GPT Realtime Translate

OpenAI realtime model for streaming speech-to-speech translation

GPT Realtime Translate focuses on low-latency speech translation for cross-language calls, meetings and voice products.

OpenAIrealtime

GPT Realtime Whisper

OpenAI streaming speech-to-text model for realtime transcription

GPT Realtime Whisper is an OpenAI speech-to-text model for transcription, captions, voice input and audio content processing.

OpenAItranscribe

GPT-4o Transcribe

Speech-to-text model powered by GPT-4o

GPT-4o Transcribe is an OpenAI speech-to-text model for transcription, captions, voice input and audio content processing.

OpenAItranscribe

GPT-4o mini Transcribe

Cost-efficient speech-to-text model powered by GPT-4o mini

GPT-4o mini Transcribe is an OpenAI speech-to-text model for transcription, captions, voice input and audio content processing.

OpenAItts

GPT-4o mini TTS

Text-to-speech model powered by GPT-4o mini

GPT-4o mini TTS turns text into speech for narration, assistant replies, read-aloud experiences and voice-enabled products.

domainAnthropic24 models

Anthropicclaude

Claude Opus 5

Claude Opus 5 for complex coding, long-running agents, professional knowledge work and scientific research

Claude Opus 5 is an Anthropic Claude model for complex coding, long-running agents, professional knowledge work and scientific research, with a focus on long-context understanding, writing quality and reliable analysis.

1000K contextMultimodalInput $5/1M tokensOutput $25/1M tokensReleased July 2026
Anthropicclaude

Claude Sonnet 5

Claude Sonnet 5 for coding, agents and professional work at scale

Claude Sonnet 5 is an Anthropic Claude model for coding, agents and professional work at scale, with a focus on long-context understanding, writing quality and reliable analysis.

1000K contextMultimodalInput $2/1M tokensOutput $10/1M tokensReleased June 2026
Anthropicclaude

Claude Fable 5

Claude Fable 5 for ambitious knowledge work, complex coding, long-running agents and multi-day enterprise workflows

Claude Fable 5 is an Anthropic Claude model for ambitious knowledge work, complex coding, long-running agents and multi-day enterprise workflows, with a focus on long-context understanding, writing quality and reliable analysis.

1000K contextMultimodalInput $10/1M tokensOutput $50/1M tokensReleased June 2026
Anthropicclaude

Claude Mythos 5

Claude Mythos 5 for cybersecurity, biology research, healthcare benchmarks and trusted high-risk research workflows

Claude Mythos 5 is an Anthropic Claude model for cybersecurity, biology research, healthcare benchmarks and trusted high-risk research workflows, with a focus on long-context understanding, writing quality and reliable analysis.

Released June 2026
Anthropicclaude

Claude Opus 4.8

Claude Opus 4.8 for serious coding, agentic workflows, long-running tasks and high-stakes professional work

Claude Opus 4.8 is an Anthropic Claude model for serious coding, agentic workflows, long-running tasks and high-stakes professional work, with a focus on long-context understanding, writing quality and reliable analysis.

1000K contextMultimodalInput $5/1M tokensOutput $25/1M tokensReleased May 2026
Anthropicclaude-opus

Claude Opus 4.7

Claude Opus 4.7: a claude-opus model from Anthropic, ~1M context, knowledge cutoff 2026-01-31

Claude Opus 4.7 is a claude-opus model from Anthropic (~1M context, input around $5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1000K contextMultimodalInput $5/1M tokensOutput $25/1M tokensReleased April 2026
Anthropicclaude-sonnet

Claude Sonnet 4.6

Claude Sonnet 4.6: a claude-sonnet model from Anthropic, ~1M context, knowledge cutoff 2025-08-31

Claude Sonnet 4.6 is a claude-sonnet model from Anthropic (~1M context, input around $3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1000K contextMultimodalInput $3/1M tokensOutput $15/1M tokensReleased February 2026
Anthropicclaude-opus

Claude Opus 4.6

Claude Opus 4.6: a claude-opus model from Anthropic, ~1M context, knowledge cutoff 2025-05-31

Claude Opus 4.6 is a claude-opus model from Anthropic (~1M context, input around $5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1000K contextMultimodalInput $5/1M tokensOutput $25/1M tokensReleased February 2026
Anthropicclaude-opus

Claude Opus 4.5

Claude Opus 4.5: a claude-opus model from Anthropic, ~200K context, knowledge cutoff 2025-05

Claude Opus 4.5 is a claude-opus model from Anthropic (~200K context, input around $5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $5/1M tokensOutput $25/1M tokensReleased November 2025
Anthropicclaude-haiku

Claude Haiku 4.5 (latest)

Claude Haiku 4.5 (latest): a claude-haiku model from Anthropic, ~200K context, knowledge cutoff 2025-02-28

Claude Haiku 4.5 (latest) is a claude-haiku model from Anthropic (~200K context, input around $1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $1/1M tokensOutput $5/1M tokensReleased October 2025
Anthropicclaude

Claude Sonnet 4.5

Claude Sonnet model optimized for coding, agents and complex workflows

Claude Sonnet 4.5 is positioned as a high-quality model for coding, long-form reasoning, agent workflows and structured professional writing. It is a strong option when reliability and instruction following matter.

1000K contextMultimodalInput $3/1M tokensOutput $15/1M tokensReleased September 2025
Anthropicclaude-opus

Claude Opus 4.1

Claude Opus 4.1: a claude-opus model from Anthropic, ~200K context, knowledge cutoff 2025-03-31

Claude Opus 4.1 is a claude-opus model from Anthropic (~200K context, input around $15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextMultimodalInput $15/1M tokensOutput $75/1M tokensReleased August 2025
Anthropicclaude-opus

Claude Opus 4 (latest)

Claude Opus 4 (latest): a claude-opus model from Anthropic, ~200K context, knowledge cutoff 2025-03-31

Claude Opus 4 (latest) is a claude-opus model from Anthropic (~200K context, input around $15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released May 2025
Anthropicclaude-sonnet

Claude Sonnet 4 (latest)

Claude Sonnet 4 (latest): a claude-sonnet model from Anthropic, ~200K context, knowledge cutoff 2025-03-31

Claude Sonnet 4 (latest) is a claude-sonnet model from Anthropic (~200K context, input around $3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released May 2025
Anthropicclaude

Claude Sonnet 4

Claude Sonnet 4 model profile for capabilities, pricing and use cases

Claude Sonnet 4 is a large language model from Anthropic. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

Released May 2025
Anthropicclaude

Claude Opus 4

High-end Claude model for difficult reasoning, coding and long-running work

Claude Opus 4 is positioned for demanding tasks that need stronger reasoning, deeper code understanding and careful execution over longer workflows.

Released May 2025
Anthropicclaude-sonnet

Claude Sonnet 3.7

Claude Sonnet 3.7: a claude-sonnet model from Anthropic, ~200K context, knowledge cutoff 2024-10-31

Claude Sonnet 3.7 is a claude-sonnet model from Anthropic (~200K context, input around $3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released February 2025
Anthropicclaude-haiku

Claude Haiku 3.5 (latest)

Claude Haiku 3.5 (latest): a claude-haiku model from Anthropic, ~200K context, knowledge cutoff 2024-07-31

Claude Haiku 3.5 (latest) is a claude-haiku model from Anthropic (~200K context, input around $0.8/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released October 2024
Anthropicclaude-haiku

Claude Haiku 3.5

Claude Haiku 3.5: a claude-haiku model from Anthropic, ~200K context, knowledge cutoff 2024-07-31

Claude Haiku 3.5 is a claude-haiku model from Anthropic (~200K context, input around $0.8/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released October 2024
Anthropicclaude-sonnet

Claude Sonnet 3.5 v2

Claude Sonnet 3.5 v2: a claude-sonnet model from Anthropic, ~200K context, knowledge cutoff 2024-04-30

Claude Sonnet 3.5 v2 is a claude-sonnet model from Anthropic (~200K context, input around $3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released October 2024
Anthropicclaude-haiku

Claude Haiku 3

Claude Haiku 3: a claude-haiku model from Anthropic, ~200K context, knowledge cutoff 2023-08-31

Claude Haiku 3 is a claude-haiku model from Anthropic (~200K context, input around $0.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released March 2024
Anthropicclaude-sonnet

Claude Sonnet 3

Claude Sonnet 3: a claude-sonnet model from Anthropic, ~200K context, knowledge cutoff 2023-08-31

Claude Sonnet 3 is a claude-sonnet model from Anthropic (~200K context, input around $3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released March 2024
Anthropicclaude-opus

Claude Opus 3

Claude Opus 3: a claude-opus model from Anthropic, ~200K context, knowledge cutoff 2023-08-31

Claude Opus 3 is a claude-opus model from Anthropic (~200K context, input around $15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released February 2024
Anthropicclaude

Claude Haiku 4

Claude Haiku 4 model profile for capabilities, pricing and use cases

Claude Haiku 4 is a large language model from Anthropic. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

domainGoogle43 models

Googlegemini-flash-lite

Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite: a gemini-flash-lite model from Google, ~1M context, knowledge cutoff 2026-03

Gemini 3.5 Flash Lite is a gemini-flash-lite model from Google (~1M context, input around $0.3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.3/1M tokensOutput $2.5/1M tokensReleased July 2026
Googlegemini-flash

Gemini 3.6 Flash

Gemini 3.6 Flash: a gemini-flash model from Google, ~1M context, knowledge cutoff 2026-03

Gemini 3.6 Flash is a gemini-flash model from Google (~1M context, input around $1.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $1.5/1M tokensOutput $7.5/1M tokensReleased July 2026
Googlegemini

Gemini Omni Flash Preview

Gemini Omni Flash Preview: a gemini model from Google, ~131.1K context

Gemini Omni Flash Preview is a gemini model from Google (~131.1K context, input around $1.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $1.5/1M tokensOutput $17.5/1M tokensReleased June 2026
Googlegemini-flash-lite

Nano Banana 2 Lite

Nano Banana 2 Lite: a gemini-flash-lite model from Google, ~65.5K context, knowledge cutoff 2025-01

Nano Banana 2 Lite is a gemini-flash-lite model from Google (~65.5K context, input around $0.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

66K contextMultimodalInput $0.25/1M tokensOutput $30/1M tokensReleased June 2026
Googlegemini-pro

Gemini 3.5 Live Translate Preview

Gemini 3.5 Live Translate Preview: a gemini-pro model from Google, ~16.4K context, knowledge cutoff 2025-01

Gemini 3.5 Live Translate Preview is a gemini-pro model from Google (~16.4K context, input around $3.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

16K contextInput $3.5/1M tokensOutput $21/1M tokensReleased June 2026
Googlegemini-flash

Nano Banana 2

Nano Banana 2: a gemini-flash model from Google, ~65.5K context, knowledge cutoff 2025-01

Nano Banana 2 is a gemini-flash model from Google (~65.5K context, input around $0.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

66K contextMultimodalInput $0.5/1M tokensOutput $60/1M tokensReleased May 2026
Googlegemini-pro

Nano Banana Pro

Nano Banana Pro: a gemini-pro model from Google, ~131.1K context, knowledge cutoff 2025-01

Nano Banana Pro is a gemini-pro model from Google (~131.1K context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $2/1M tokensOutput $120/1M tokensReleased May 2026
Googlegemini-flash

Gemini Flash Latest

Gemini Flash Latest: a gemini-flash model from Google, ~1M context, knowledge cutoff 2025-01

Gemini Flash Latest is a gemini-flash model from Google (~1M context, input around $1.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $1.5/1M tokensOutput $9/1M tokensReleased May 2026
Googlegemini-flash

Gemini 3.5 Flash

Gemini 3.5 Flash: a gemini-flash model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3.5 Flash is a gemini-flash model from Google (~1M context, input around $1.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $1.5/1M tokensOutput $9/1M tokensReleased May 2026
Googlegemini-flash-lite

Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite: a gemini-flash-lite model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3.1 Flash Lite is a gemini-flash-lite model from Google (~1M context, input around $0.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.25/1M tokensOutput $1.5/1M tokensReleased May 2026
Googlegemini-flash-lite

Gemini Flash-Lite Latest

Gemini Flash-Lite Latest: a gemini-flash-lite model from Google, ~1M context, knowledge cutoff 2025-01

Gemini Flash-Lite Latest is a gemini-flash-lite model from Google (~1M context, input around $0.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.25/1M tokensOutput $1.5/1M tokensReleased May 2026
Googlegemini

Gemini Embedding 2

Gemini Embedding 2: a gemini model from Google, ~8.2K context, knowledge cutoff 2025-11

Gemini Embedding 2 is a gemini model from Google (~8.2K context, input around $0.2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextMultimodalInput $0.2/1M tokensOutput $0/1M tokensReleased April 2026
Googlegemini-pro

Deep Research Preview (Apr-21-2026)

Deep Research Preview (Apr-21-2026): a gemini-pro model from Google, ~131.1K context, knowledge cutoff 2025-01

Deep Research Preview (Apr-21-2026) is a gemini-pro model from Google (~131.1K context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $2/1M tokensOutput $12/1M tokensReleased April 2026
Googlegemini-pro

Deep Research Max Preview (Apr-21-2026)

Deep Research Max Preview (Apr-21-2026): a gemini-pro model from Google, ~131.1K context, knowledge cutoff 2025-01

Deep Research Max Preview (Apr-21-2026) is a gemini-pro model from Google (~131.1K context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $2/1M tokensOutput $12/1M tokensReleased April 2026
Googlegemini-flash

Gemini 3.1 Flash TTS Preview

Gemini 3.1 Flash TTS Preview: a gemini-flash model from Google, ~8.2K context, knowledge cutoff 2025-01

Gemini 3.1 Flash TTS Preview is a gemini-flash model from Google (~8.2K context, input around $1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $1/1M tokensOutput $20/1M tokensReleased April 2026
Googlegemini

Gemini Robotics-ER 1.6 Preview

Gemini Robotics-ER 1.6 Preview: a gemini model from Google, ~131.1K context, knowledge cutoff 2025-01

Gemini Robotics-ER 1.6 Preview is a gemini model from Google (~131.1K context, input around $1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $1/1M tokensOutput $5/1M tokensReleased April 2026
Googlegemma

Gemma 4 31B IT

Gemma 4 31B IT: a gemma model from Google, ~262.1K context

Gemma 4 31B IT is a gemma model from Google (~262.1K context). Suitable for assistants, content generation, knowledge Q&A and business automation

Released April 2026
Googlegemma

Gemma 4 26B A4B IT

Gemma 4 26B A4B IT: a gemma model from Google, ~262.1K context

Gemma 4 26B A4B IT is a gemma model from Google (~262.1K context). Suitable for assistants, content generation, knowledge Q&A and business automation

Released April 2026
Googleveo

Veo 3.1 lite

Veo 3.1 lite: a veo model from Google

Veo 3.1 lite is a veo model from Google. Suitable for assistants, content generation, knowledge Q&A and business automation

Released March 2026
Googlelyria

Lyria 3 Pro Preview

Lyria 3 Pro Preview: a lyria model from Google, ~1M context

Lyria 3 Pro Preview is a lyria model from Google (~1M context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0/1M tokensOutput $0/1M tokensReleased March 2026
Googlelyria

Lyria 3 Clip Preview

Lyria 3 Clip Preview: a lyria model from Google, ~1M context

Lyria 3 Clip Preview is a lyria model from Google (~1M context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0/1M tokensOutput $0/1M tokensReleased March 2026
Googlegemini-flash-lite

Gemini 3.1 Flash Lite Preview

Gemini 3.1 Flash Lite Preview: a gemini-flash-lite model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3.1 Flash Lite Preview is a gemini-flash-lite model from Google (~1M context, input around $0.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.25/1M tokensOutput $1.5/1M tokensReleased March 2026
Googlegemini-flash

Nano Banana 2

Nano Banana 2: a gemini-flash model from Google, ~65.5K context, knowledge cutoff 2025-01

Nano Banana 2 is a gemini-flash model from Google (~65.5K context, input around $0.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

66K contextMultimodalInput $0.5/1M tokensOutput $60/1M tokensReleased February 2026
Googlegemini-pro

Gemini 3.1 Pro Preview Custom Tools

Gemini 3.1 Pro Preview Custom Tools: a gemini-pro model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3.1 Pro Preview Custom Tools is a gemini-pro model from Google (~1M context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $2/1M tokensOutput $12/1M tokensReleased February 2026
Googlegemini-pro

Gemini 3.1 Pro Preview

Gemini 3.1 Pro Preview: a gemini-pro model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3.1 Pro Preview is a gemini-pro model from Google (~1M context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $2/1M tokensOutput $12/1M tokensReleased February 2026
Googlegemini-flash

Gemini 3 Flash Preview

Gemini 3 Flash Preview: a gemini-flash model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3 Flash Preview is a gemini-flash model from Google (~1M context, input around $0.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.5/1M tokensOutput $3/1M tokensReleased December 2025
Googlegemini-pro

Nano Banana Pro

Nano Banana Pro: a gemini-pro model from Google, ~131.1K context, knowledge cutoff 2025-01

Nano Banana Pro is a gemini-pro model from Google (~131.1K context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $2/1M tokensOutput $120/1M tokensReleased November 2025
Googlegemini-pro

Gemini 3 Pro Preview

Gemini 3 Pro Preview: a gemini-pro model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 3 Pro Preview is a gemini-pro model from Google (~1M context, input around $2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $2/1M tokensOutput $12/1M tokensReleased November 2025
Googleveo

Veo 3.1

Veo 3.1: a veo model from Google

Veo 3.1 is a veo model from Google. Suitable for assistants, content generation, knowledge Q&A and business automation

Released October 2025
Googleveo

Veo 3.1 fast

Veo 3.1 fast: a veo model from Google

Veo 3.1 fast is a veo model from Google. Suitable for assistants, content generation, knowledge Q&A and business automation

Released October 2025
Googlegemini-pro

Gemini 2.5 Computer Use Preview 10-2025

Gemini 2.5 Computer Use Preview 10-2025: a gemini-pro model from Google, ~131.1K context, knowledge cutoff 2025-01

Gemini 2.5 Computer Use Preview 10-2025 is a gemini-pro model from Google (~131.1K context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextMultimodalInput $1.25/1M tokensOutput $10/1M tokensReleased October 2025
Googlegemini-flash

Nano Banana

Nano Banana: a gemini-flash model from Google, ~32.8K context, knowledge cutoff 2024-06

Nano Banana is a gemini-flash model from Google (~32.8K context, input around $0.3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

33K contextMultimodalInput $0.3/1M tokensOutput $30/1M tokensReleased August 2025
Googlegemini-pro

Gemini 2.5 Pro

Gemini 2.5 Pro: a gemini-pro model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 2.5 Pro is a gemini-pro model from Google (~1M context, input around $1.25/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $1.25/1M tokensOutput $10/1M tokensReleased June 2025
Googlegemini-flash-lite

Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite: a gemini-flash-lite model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 2.5 Flash-Lite is a gemini-flash-lite model from Google (~1M context, input around $0.1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.1/1M tokensOutput $0.4/1M tokensReleased June 2025
Googlegemini-flash

Gemini 2.5 Flash

Gemini 2.5 Flash: a gemini-flash model from Google, ~1M context, knowledge cutoff 2025-01

Gemini 2.5 Flash is a gemini-flash model from Google (~1M context, input around $0.3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.3/1M tokensOutput $2.5/1M tokensReleased June 2025
Googlegemini

Gemini Embedding 001

Gemini Embedding 001: a gemini model from Google, ~2K context, knowledge cutoff 2025-05

Gemini Embedding 001 is a gemini model from Google (~2K context, input around $0.15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

2K contextInput $0.15/1M tokensOutput $0/1M tokensReleased May 2025
Googlegemini-flash

Gemini 2.5 Pro Preview TTS

Gemini 2.5 Pro Preview TTS: a gemini-flash model from Google, ~8.2K context, knowledge cutoff 2025-01

Gemini 2.5 Pro Preview TTS is a gemini-flash model from Google (~8.2K context, input around $1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $1/1M tokensOutput $20/1M tokensReleased May 2025
Googlegemini-flash

Gemini 2.5 Flash Preview TTS

Gemini 2.5 Flash Preview TTS: a gemini-flash model from Google, ~8.2K context, knowledge cutoff 2025-01

Gemini 2.5 Flash Preview TTS is a gemini-flash model from Google (~8.2K context, input around $0.5/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

8K contextInput $0.5/1M tokensOutput $10/1M tokensReleased May 2025
Googlegemini-pro

Gemini 2.5 Pro

Google advanced multimodal model for long-context and reasoning-heavy tasks

Gemini 2.5 Pro is Google's high-capability Gemini model, useful for complex reasoning, long-context analysis and multimodal applications across text and media inputs.

1049K contextMultimodalInput $1.25/1M tokensOutput $10/1M tokensReleased March 2025
Googlegemini-flash

Gemini 2.5 Flash

Fast Gemini model for low-latency multimodal and high-throughput tasks

Gemini 2.5 Flash is optimized for speed and efficiency, making it suitable for interactive products, lightweight reasoning and high-volume calls.

1049K contextMultimodalInput $0.3/1M tokensOutput $2.5/1M tokensReleased March 2025
Googlegemini-flash-lite

Gemini 2.0 Flash-Lite

Gemini 2.0 Flash-Lite: a gemini-flash-lite model from Google, ~1M context, knowledge cutoff 2024-06

Gemini 2.0 Flash-Lite is a gemini-flash-lite model from Google (~1M context, input around $0.075/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.075/1M tokensOutput $0.3/1M tokensReleased December 2024
Googlegemini-flash

Gemini 2.0 Flash

Gemini 2.0 Flash: a gemini-flash model from Google, ~1M context, knowledge cutoff 2024-06

Gemini 2.0 Flash is a gemini-flash model from Google (~1M context, input around $0.1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1049K contextMultimodalInput $0.1/1M tokensOutput $0.4/1M tokensReleased December 2024
Googlegemini

Gemini Experimental 1206

Gemini Experimental 1206 model profile for capabilities, pricing and use cases

Gemini Experimental 1206 is a large language model from Google. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

domainDeepSeek6 models

DeepSeekdeepseek-thinking

DeepSeek Reasoner

DeepSeek Reasoner: a deepseek-thinking model from DeepSeek, ~1M context, knowledge cutoff 2025-09

DeepSeek Reasoner is a deepseek-thinking model from DeepSeek (~1M context, input around $0.14/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

1000K contextInput $0.14/1M tokensOutput $0.28/1M tokensReleased December 2025
DeepSeekdeepseek

DeepSeek Chat

DeepSeek Chat model profile for capabilities, pricing and use cases

DeepSeek Chat is a large language model from DeepSeek. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

1000K contextInput $0.14/1M tokensOutput $0.28/1M tokensReleased December 2025
DeepSeekdeepseek

DeepSeek V4 Pro

DeepSeek V4 Pro for higher-quality Chinese tasks, coding and enterprise workloads

DeepSeek V4 Pro is a DeepSeek model for higher-quality Chinese tasks, coding and enterprise workloads, often evaluated for Chinese tasks, coding and cost efficiency.

1000K contextInput ¥3/1M tokensOutput ¥6/1M tokens
DeepSeekdeepseek

DeepSeek V4 Flash

DeepSeek V4 Flash for low-latency chat, coding and high-volume production workloads

DeepSeek V4 Flash is a DeepSeek model for low-latency chat, coding and high-volume production workloads, often evaluated for Chinese tasks, coding and cost efficiency.

1000K contextInput ¥1/1M tokensOutput ¥2/1M tokens
DeepSeekdeepseek

DeepSeek V3

DeepSeek general chat model with strong Chinese, coding and value performance

DeepSeek V3 is a general-purpose chat model known for strong value, Chinese language tasks and coding assistance. It is often considered for cost-sensitive production workloads.

66K contextInput ¥2/1M tokensOutput ¥8/1M tokens
DeepSeekdeepseek

DeepSeek R1

DeepSeek reasoning model for logic, math and code analysis

DeepSeek R1 focuses on reasoning-heavy tasks such as mathematical thinking, multi-step logic and code analysis while remaining attractive for value-conscious teams.

131K contextInput ¥4/1M tokensOutput ¥16/1M tokens

domainAlibaba Cloud (Qwen)5 models

Alibaba Cloud (Qwen)qwen

Qwen3.7 Max

Qwen3.7 Max for advanced reasoning, coding, high-quality Chinese tasks and enterprise workloads

Qwen3.7 Max is a Qwen model from Alibaba Cloud for advanced reasoning, coding, high-quality Chinese tasks and enterprise workloads, commonly evaluated for Chinese enterprise and automation workloads.

1000K contextInput ¥12/1M tokensOutput ¥36/1M tokens
Alibaba Cloud (Qwen)qwen

Qwen3.7 Plus

Qwen3.7 Plus for balanced enterprise Q&A, writing, coding and agent workflows

Qwen3.7 Plus is a Qwen model from Alibaba Cloud for balanced enterprise Q&A, writing, coding and agent workflows, commonly evaluated for Chinese enterprise and automation workloads.

1000K contextInput ¥2/1M tokensOutput ¥8/1M tokens
Alibaba Cloud (Qwen)qwen

Qwen Turbo

Qwen Turbo model profile for capabilities, pricing and use cases

Qwen Turbo is a large language model from Qwen. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

131K contextInput ¥0.3/1M tokensOutput ¥0.6/1M tokens
Alibaba Cloud (Qwen)qwen

Qwen Plus

Balanced Qwen model for enterprise assistants, writing and coding support

Qwen Plus is a balanced model in the Qwen family, suitable for common enterprise workloads, Chinese writing, knowledge Q&A and code assistance.

1000K contextInput ¥0.8/1M tokensOutput ¥2/1M tokens
Alibaba Cloud (Qwen)qwen

Qwen Max

Qwen high-capability model for advanced Chinese and coding tasks

Qwen Max is a high-end model in the Qwen family. It is useful for Chinese enterprise scenarios, complex writing, knowledge workflows and code assistance.

33K contextInput ¥2.4/1M tokensOutput ¥9.6/1M tokens

domainMoonshot (Kimi)16 models

Moonshot (Kimi)kimi-k3

Kimi K3

Kimi K3 for long-horizon coding, knowledge work, visual understanding and deep reasoning

Kimi K3 is Moonshot AI's flagship model with a 1M-token context window and native visual understanding for long-horizon coding, knowledge work and deep reasoning.

1049K contextMultimodalInput ¥20/1M tokensOutput ¥100/1M tokensReleased July 2026
Moonshot (Kimi)kimi-k2

Kimi K2.7 Code

Kimi K2.7 Code: a kimi-k2 model from Moonshot AI, ~262.1K context, knowledge cutoff 2025-01

Kimi K2.7 Code is a kimi-k2 model from Moonshot AI (~262.1K context, input around $0.95/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextMultimodalInput $0.95/1M tokensOutput $4/1M tokensReleased June 2026
Moonshot (Kimi)kimi-k2

Kimi K2.7 Code HighSpeed

Kimi K2.7 Code HighSpeed: a kimi-k2 model from Moonshot AI, ~262.1K context, knowledge cutoff 2025-01

Kimi K2.7 Code HighSpeed is a kimi-k2 model from Moonshot AI (~262.1K context, input around $1.9/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextMultimodalInput $1.9/1M tokensOutput $8/1M tokensReleased June 2026
Moonshot (Kimi)kimi-k2

Kimi K2.6

Kimi K2.6: a kimi-k2 model from Moonshot AI, ~262.1K context, knowledge cutoff 2025-01

Kimi K2.6 is a kimi-k2 model from Moonshot AI (~262.1K context, input around $0.95/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextMultimodalInput $0.95/1M tokensOutput $4/1M tokensReleased April 2026
Moonshot (Kimi)kimi-k2

Kimi K2.5

Kimi K2.5: a kimi-k2 model from Moonshot AI, ~262.1K context, knowledge cutoff 2025-01

Kimi K2.5 is a kimi-k2 model from Moonshot AI (~262.1K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextMultimodalInput $0.6/1M tokensOutput $3/1M tokensReleased January 2026
Moonshot (Kimi)kimi-thinking

Kimi K2 Thinking

Kimi K2 Thinking: a kimi-thinking model from Moonshot AI, ~262.1K context, knowledge cutoff 2024-08

Kimi K2 Thinking is a kimi-thinking model from Moonshot AI (~262.1K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextInput $0.6/1M tokensOutput $2.5/1M tokensReleased November 2025
Moonshot (Kimi)kimi-thinking

Kimi K2 Thinking Turbo

Kimi K2 Thinking Turbo: a kimi-thinking model from Moonshot AI, ~262.1K context, knowledge cutoff 2024-08

Kimi K2 Thinking Turbo is a kimi-thinking model from Moonshot AI (~262.1K context, input around $1.15/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextInput $1.15/1M tokensOutput $8/1M tokensReleased November 2025
Moonshot (Kimi)kimi-k2

Kimi K2 0905

Kimi K2 0905: a kimi-k2 model from Moonshot AI, ~262.1K context, knowledge cutoff 2024-10

Kimi K2 0905 is a kimi-k2 model from Moonshot AI (~262.1K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextInput $0.6/1M tokensOutput $2.5/1M tokensReleased September 2025
Moonshot (Kimi)kimi-k2

Kimi K2 Turbo

Kimi K2 Turbo: a kimi-k2 model from Moonshot AI, ~262.1K context, knowledge cutoff 2024-10

Kimi K2 Turbo is a kimi-k2 model from Moonshot AI (~262.1K context, input around $2.4/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

262K contextInput $2.4/1M tokensOutput $10/1M tokensReleased September 2025
Moonshot (Kimi)kimi-k2

Kimi K2 0711

Kimi K2 0711: a kimi-k2 model from Moonshot AI, ~131.1K context, knowledge cutoff 2024-10

Kimi K2 0711 is a kimi-k2 model from Moonshot AI (~131.1K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextInput $0.6/1M tokensOutput $2.5/1M tokensReleased July 2025
Moonshot (Kimi)kimi

Moonshot v1 128K

Moonshot v1 128K for very long-document analysis, retrieval-augmented reading and complex context workflows

Moonshot v1 128K is a Moonshot/Kimi model for very long-document analysis, retrieval-augmented reading and complex context workflows, often evaluated for Chinese document and knowledge workflows.

131K contextInput ¥10/1M tokensOutput ¥30/1M tokens
Moonshot (Kimi)kimi-k2

Kimi K2

Kimi K2 for complex reasoning, coding, agent workflows and higher-value production tasks

Kimi K2 is a Moonshot/Kimi model for complex reasoning, coding, agent workflows and higher-value production tasks, often evaluated for Chinese document and knowledge workflows.

Moonshot (Kimi)kimi-k2

Kimi K2.6

Kimi K2.6 for complex reasoning, coding, agent workflows and production assistants

Kimi K2.6 is a Moonshot/Kimi model for complex reasoning, coding, agent workflows and production assistants, often evaluated for Chinese document and knowledge workflows.

262K contextMultimodalInput $0.95/1M tokensOutput $4/1M tokens
Moonshot (Kimi)kimi-k2

Kimi K2.5

Kimi K2.5 for coding, long-running tasks and agent workflows

Kimi K2.5 is a Moonshot/Kimi model for coding, long-running tasks and agent workflows, often evaluated for Chinese document and knowledge workflows.

262K contextMultimodalInput $0.6/1M tokensOutput $3/1M tokens
Moonshot (Kimi)kimi

Moonshot v1 8K

Moonshot v1 8K model profile for capabilities, pricing and use cases

Moonshot v1 8K is a large language model from Moonshot AI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

8K contextInput ¥2/1M tokensOutput ¥10/1M tokens
Moonshot (Kimi)kimi

Moonshot v1 32K

Moonshot v1 32K model profile for capabilities, pricing and use cases

Moonshot v1 32K is a large language model from Moonshot AI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

33K contextInput ¥5/1M tokensOutput ¥20/1M tokens

domainZhipu AI18 models

Zhipu AIglm-flash

GLM-4.7-FlashX

GLM-4.7-FlashX: a glm-flash model from Zhipu AI, ~200K context, knowledge cutoff 2025-04

GLM-4.7-FlashX is a glm-flash model from Zhipu AI (~200K context, input around $0.07/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextInput $0.07/1M tokensOutput $0.4/1M tokensReleased January 2026
Zhipu AIglm-flash

GLM-4.7-Flash

GLM-4.7-Flash: a glm-flash model from Zhipu AI, ~200K context, knowledge cutoff 2025-04

GLM-4.7-Flash is a glm-flash model from Zhipu AI (~200K context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

200K contextInput $0/1M tokensOutput $0/1M tokensReleased January 2026
Zhipu AIglm

GLM-4.7

GLM-4.7: a glm model from Zhipu AI, ~204.8K context, knowledge cutoff 2025-04

GLM-4.7 is a glm model from Zhipu AI (~204.8K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

205K contextInput $0.6/1M tokensOutput $2.2/1M tokensReleased December 2025
Zhipu AIglm

GLM-4.6V

GLM-4.6V: a glm model from Zhipu AI, ~128K context, knowledge cutoff 2025-04

GLM-4.6V is a glm model from Zhipu AI (~128K context, input around $0.3/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

128K contextMultimodalInput $0.3/1M tokensOutput $0.9/1M tokensReleased December 2025
Zhipu AIglm

GLM-4.6

GLM-4.6: a glm model from Zhipu AI, ~204.8K context, knowledge cutoff 2025-04

GLM-4.6 is a glm model from Zhipu AI (~204.8K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

205K contextInput $0.6/1M tokensOutput $2.2/1M tokensReleased September 2025
Zhipu AIglm

GLM-4.5V

GLM-4.5V: a glm model from Zhipu AI, ~64K context, knowledge cutoff 2025-04

GLM-4.5V is a glm model from Zhipu AI (~64K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

64K contextMultimodalInput $0.6/1M tokensOutput $1.8/1M tokensReleased August 2025
Zhipu AIglm-air

GLM-4.5-Air

GLM-4.5-Air: a glm-air model from Zhipu AI, ~131.1K context, knowledge cutoff 2025-04

GLM-4.5-Air is a glm-air model from Zhipu AI (~131.1K context, input around $0.2/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextInput $0.2/1M tokensOutput $1.1/1M tokensReleased July 2025
Zhipu AIglm

GLM-4.5

GLM-4.5: a glm model from Zhipu AI, ~131.1K context, knowledge cutoff 2025-04

GLM-4.5 is a glm model from Zhipu AI (~131.1K context, input around $0.6/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextInput $0.6/1M tokensOutput $2.2/1M tokensReleased July 2025
Zhipu AIglm-flash

GLM-4.5-Flash

GLM-4.5-Flash: a glm-flash model from Zhipu AI, ~131.1K context, knowledge cutoff 2025-04

GLM-4.5-Flash is a glm-flash model from Zhipu AI (~131.1K context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

131K contextInput $0/1M tokensOutput $0/1M tokensReleased July 2025
Zhipu AIglm

GLM-4-Air

GLM-4-Air for low-cost high-concurrency Chinese assistant workloads

GLM-4-Air is a Zhipu AI GLM model for low-cost high-concurrency Chinese assistant workloads, commonly evaluated for Chinese enterprise and agent workflows.

Zhipu AIglm

GLM-4-Flash

GLM-4-Flash for fast responses, lightweight automation and high-frequency conversations

GLM-4-Flash is a Zhipu AI GLM model for fast responses, lightweight automation and high-frequency conversations, commonly evaluated for Chinese enterprise and agent workflows.

Zhipu AIglm-4.5

GLM-4.5

GLM-4.5 for agent workflows, coding and complex reasoning tasks

GLM-4.5 is a Zhipu AI GLM model for agent workflows, coding and complex reasoning tasks, commonly evaluated for Chinese enterprise and agent workflows.

131K contextInput $0.6/1M tokensOutput $2.2/1M tokens
Zhipu AIglm-z1

GLM-Z1

GLM-Z1 for multi-step reasoning, math, coding and complex problem solving

GLM-Z1 is a Zhipu AI GLM model for multi-step reasoning, math, coding and complex problem solving, commonly evaluated for Chinese enterprise and agent workflows.

Zhipu AIglm-5

GLM-5.1

GLM-5.1 for advanced reasoning, coding and agent workflows

GLM-5.1 is a Zhipu AI GLM model for advanced reasoning, coding and agent workflows, commonly evaluated for Chinese enterprise and agent workflows.

200K contextInput $1.4/1M tokensOutput $4.4/1M tokens
Zhipu AIglm-5

GLM-5-Turbo

GLM-5-Turbo for cost-efficient Chinese business tasks and high-volume assistants

GLM-5-Turbo is a Zhipu AI GLM model for cost-efficient Chinese business tasks and high-volume assistants, commonly evaluated for Chinese enterprise and agent workflows.

200K contextInput $1.2/1M tokensOutput $4/1M tokens
Zhipu AIglm-5

GLM-5

GLM-5 for Chinese dialogue, coding and enterprise applications

GLM-5 is a Zhipu AI GLM model for Chinese dialogue, coding and enterprise applications, commonly evaluated for Chinese enterprise and agent workflows.

205K contextInput $1/1M tokensOutput $3.2/1M tokens
Zhipu AIglm

GLM-4-Plus

GLM-4-Plus model profile for capabilities, pricing and use cases

GLM-4-Plus is a large language model from Zhipu AI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

128K contextInput ¥5/1M tokensOutput ¥5/1M tokens
Zhipu AIglm

GLM-4

GLM-4 model profile for capabilities, pricing and use cases

GLM-4 is a large language model from Zhipu AI. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

domainMiniMax8 models

MiniMaxabab

ABAB6.5s Chat

ABAB6.5s Chat for Chinese chat, writing and business assistant workloads

ABAB6.5s Chat is a MiniMax model for Chinese chat, writing and business assistant workloads, often evaluated for Chinese assistants, generation and multimedia workflows.

MiniMaxabab

ABAB6.5 Chat

ABAB6.5 Chat for general conversation, long-text understanding and complex interactions

ABAB6.5 Chat is a MiniMax model for general conversation, long-text understanding and complex interactions, often evaluated for Chinese assistants, generation and multimedia workflows.

MiniMaxminimax-text

MiniMax Text 01

MiniMax Text 01 for general language understanding, generation and agent workflows

MiniMax Text 01 is a MiniMax model for general language understanding, generation and agent workflows, often evaluated for Chinese assistants, generation and multimedia workflows.

MiniMaxminimax-m

MiniMax M1

MiniMax M1 for long-context reasoning, coding and complex task planning

MiniMax M1 is a MiniMax model for long-context reasoning, coding and complex task planning, often evaluated for Chinese assistants, generation and multimedia workflows.

MiniMaxminimax-speech

MiniMax Speech 01

MiniMax speech model for voice generation, conversation and multimedia content

MiniMax Speech 01 focuses on voice generation and multimedia experiences rather than general text conversation.

MiniMaxminimax-m

MiniMax M3

MiniMax M3 for coding, agent workflows and complex production tasks

MiniMax M3 is a MiniMax model for coding, agent workflows and complex production tasks, often evaluated for Chinese assistants, generation and multimedia workflows.

1000K contextMultimodalInput $0.3/1M tokensOutput $1.2/1M tokens
MiniMaxabab

abab7-chat

abab7-chat model profile for capabilities, pricing and use cases

abab7-chat is a large language model from MiniMax. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

MiniMaxabab

abab6.5s-chat

abab6.5s-chat model profile for capabilities, pricing and use cases

abab6.5s-chat is a large language model from MiniMax. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

domainStepFun9 models

StepFun

Step 3.5 Flash 2603

Step 3.5 Flash 2603: a AI model from stepfun, ~256K context, knowledge cutoff 2025-01

Step 3.5 Flash 2603 is a AI model from stepfun (~256K context, input around $0.1/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released April 2026
StepFunstep-1

Step 1 32K

Step 1 32K for document understanding, summarization and knowledge Q&A

Step 1 32K is a StepFun model for document understanding, summarization and knowledge Q&A, commonly evaluated for Chinese assistants, document and multimodal workflows.

33K contextInput $2.05/1M tokensOutput $9.59/1M tokens
StepFunstep-3

Step 3.5 Flash

Step 3.5 Flash for reasoning, Q&A and business assistant workloads

Step 3.5 Flash is a StepFun model for reasoning, Q&A and business assistant workloads, commonly evaluated for Chinese assistants, document and multimodal workflows.

StepFunstep-1

Step 1 128K

Step 1 128K for very long-document analysis and retrieval-augmented workflows

Step 1 128K is a StepFun model for very long-document analysis and retrieval-augmented workflows, commonly evaluated for Chinese assistants, document and multimodal workflows.

StepFunstep-2

Step 2 Mini

Step 2 Mini for low-latency and cost-sensitive high-volume workloads

Step 2 Mini is a StepFun model for low-latency and cost-sensitive high-volume workloads, commonly evaluated for Chinese assistants, document and multimodal workflows.

StepFunstep-vision

Step 1V 8K

Step 1V 8K for visual Q&A, multimodal analysis and image-text understanding

Step 1V 8K is a StepFun model for visual Q&A, multimodal analysis and image-text understanding, commonly evaluated for Chinese assistants, document and multimodal workflows.

StepFunstep-3

Step 3.7 Flash

Step 3.7 Flash for efficient reasoning, complex Chinese tasks and production assistants

Step 3.7 Flash is a StepFun model for efficient reasoning, complex Chinese tasks and production assistants, commonly evaluated for Chinese assistants, document and multimodal workflows.

MultimodalInput ¥1.35/1M tokensOutput ¥8.1/1M tokens
StepFunstep

Step-1 8K

Step-1 8K model profile for capabilities, pricing and use cases

Step-1 8K is a large language model from StepFun. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

StepFunstep

Step-2 16K

Step-2 16K model profile for capabilities, pricing and use cases

Step-2 16K is a large language model from StepFun. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

16K contextInput $5.21/1M tokensOutput $16.44/1M tokens

domainVolcengine7 models

Volcenginedoubao

Doubao Pro

Doubao Pro for higher-quality Chinese assistants, content generation and business automation

Doubao Pro is a Volcengine Doubao model for higher-quality Chinese assistants, content generation and business automation, often evaluated for Chinese enterprise and multimodal workflows.

Volcenginedoubao

Doubao Lite

Doubao Lite for low-cost high-concurrency chat and lightweight text tasks

Doubao Lite is a Volcengine Doubao model for low-cost high-concurrency chat and lightweight text tasks, often evaluated for Chinese enterprise and multimodal workflows.

Volcenginedoubao-seed

Doubao Seed 1.6

Doubao Seed 1.6 for general chat, reasoning and agent workflow evaluation

Doubao Seed 1.6 is a Volcengine Doubao model for general chat, reasoning and agent workflow evaluation, often evaluated for Chinese enterprise and multimodal workflows.

Input ¥0.8/1M tokensOutput ¥2/1M tokens
Volcenginedoubao-thinking

Doubao Seed 1.6 Thinking

Doubao Seed 1.6 Thinking for complex reasoning, multi-step analysis and coding assistance

Doubao Seed 1.6 Thinking is a Volcengine Doubao model for complex reasoning, multi-step analysis and coding assistance, often evaluated for Chinese enterprise and multimodal workflows.

Volcenginedoubao-vision

Doubao Vision Pro

Doubao Vision Pro for image-text analysis, multimodal Q&A and visual content understanding

Doubao Vision Pro is a Volcengine Doubao model for image-text analysis, multimodal Q&A and visual content understanding, often evaluated for Chinese enterprise and multimodal workflows.

Volcenginedoubao

Doubao Pro 32K

Doubao Pro 32K model profile for capabilities, pricing and use cases

Doubao Pro 32K is a large language model from Volcengine. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

Volcenginedoubao

Doubao Lite 32K

Doubao Lite 32K model profile for capabilities, pricing and use cases

Doubao Lite 32K is a large language model from Volcengine. This entry summarizes its positioning, typical use cases, strengths, limitations and related pricing signals for quick comparison.

domainXiaomi (MiMo)4 models

Xiaomi (MiMo)mimo

MiMo-V2.5-Pro-UltraSpeed

MiMo-V2.5-Pro-UltraSpeed: a mimo model from Xiaomi, ~1M context, knowledge cutoff 2024-12

MiMo-V2.5-Pro-UltraSpeed is a mimo model from Xiaomi (~1M context, input around $1.305/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

Released June 2026
Xiaomi (MiMo)mimo

MiMo V2.5

MiMo V2.5 for Chinese conversation, content generation and tool-use workflows

MiMo V2.5 is a Xiaomi MiMo model for Chinese conversation, content generation and tool-use workflows, useful for Chinese assistants, tool use and ecosystem-oriented applications.

1000K contextMultimodalInput ¥1/1M tokensOutput ¥2/1M tokens
Xiaomi (MiMo)mimo

MiMo V2.5 Pro

MiMo V2.5 Pro for complex reasoning, coding and long-text tasks

MiMo V2.5 Pro is a Xiaomi MiMo model for complex reasoning, coding and long-text tasks, useful for Chinese assistants, tool use and ecosystem-oriented applications.

1000K contextInput ¥3/1M tokensOutput ¥6/1M tokens
Xiaomi (MiMo)mimo

MiMo V2 Flash

MiMo V2 Flash for low-latency high-frequency chat and quick responses

MiMo V2 Flash is a Xiaomi MiMo model for low-latency high-frequency chat and quick responses, useful for Chinese assistants, tool use and ecosystem-oriented applications.

262K contextInput $0.14/1M tokensOutput $0.28/1M tokens