The Mistral-7B-Instruct-v0.2 Large Language Model (LLM) is an improved instruct fine-tuned version of Mistral-7B-Instruct-v0.1.
Replicate
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.0500
Output / 1M$0.2500
LanguageMistral 7b instruct v0.1
mistral-7b-instruct-v0.1
An instruction-tuned 7 billion parameter language model from Mistral
Replicate
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageMixtral 8x7b instruct v0.1
mistralai/mixtral-8x7b-instruct-v0.1
The Mixtral-8x7B-instruct-v0.1 Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts tuned to be a helpful assistant.
Replicate
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.3000
Output / 1M$1.0000
LanguageLlama 2 13b chat
meta/llama-2-13b-chat
A 13 billion parameter language model from Meta, fine tuned for chat completions
Replicate
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.1000
Output / 1M$0.5000
LanguageLlama 2 70b chat
meta/llama-2-70b-chat
A 70 billion parameter language model from Meta, fine tuned for chat completions
Replicate
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.6500
Output / 1M$2.7500
LanguageGPT-5.6 Sol (No Reasoning)
gpt-5.6-sol-none
GPT-5.6 Sol with reasoning disabled for fastest responses and lowest cost.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (Low Reasoning)
gpt-5.6-sol-low
GPT-5.6 Sol with low reasoning effort for lightweight thinking.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (Medium Reasoning)
gpt-5.6-sol-medium
GPT-5.6 Sol with medium reasoning effort for balanced performance.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (High Reasoning)
gpt-5.6-sol-high
GPT-5.6 Sol with high reasoning effort for complex tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (XHigh Reasoning)
gpt-5.6-sol-xhigh
GPT-5.6 Sol with xhigh reasoning effort for the hardest tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (Max Reasoning)
gpt-5.6-sol-max
GPT-5.6 Sol with max reasoning effort for the most demanding tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (No Reasoning)
gpt-5.6-terra-none
GPT-5.6 Terra with reasoning disabled for fastest responses and lowest cost.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (Low Reasoning)
gpt-5.6-terra-low
GPT-5.6 Terra with low reasoning effort for lightweight thinking.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (Medium Reasoning)
gpt-5.6-terra-medium
GPT-5.6 Terra with medium reasoning effort for balanced performance.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (High Reasoning)
gpt-5.6-terra-high
GPT-5.6 Terra with high reasoning effort for complex tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (XHigh Reasoning)
gpt-5.6-terra-xhigh
GPT-5.6 Terra with xhigh reasoning effort for the hardest tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (Max Reasoning)
gpt-5.6-terra-max
GPT-5.6 Terra with max reasoning effort for the most demanding tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (No Reasoning)
gpt-5.6-luna-none
GPT-5.6 Luna with reasoning disabled for fastest responses and lowest cost.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (Low Reasoning)
gpt-5.6-luna-low
GPT-5.6 Luna with low reasoning effort for lightweight thinking.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (Medium Reasoning)
gpt-5.6-luna-medium
GPT-5.6 Luna with medium reasoning effort for balanced performance.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (High Reasoning)
gpt-5.6-luna-high
GPT-5.6 Luna with high reasoning effort for complex tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (XHigh Reasoning)
gpt-5.6-luna-xhigh
GPT-5.6 Luna with xhigh reasoning effort for the hardest tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (Max Reasoning)
gpt-5.6-luna-max
GPT-5.6 Luna with max reasoning effort for the most demanding tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.5 (No Reasoning)
gpt-5.5-none
GPT-5.5 with reasoning disabled for fastest responses and lowest cost.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (Low Reasoning)
gpt-5.5-low
GPT-5.5 with low reasoning effort for lightweight thinking.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (Medium Reasoning)
gpt-5.5-medium
GPT-5.5 with medium reasoning effort for balanced performance.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (High Reasoning)
gpt-5.5-high
GPT-5.5 with high reasoning effort for complex tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (XHigh Reasoning)
gpt-5.5-xhigh
GPT-5.5 with xhigh reasoning effort for the hardest tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.4 (No Reasoning)
gpt-5.4-none
GPT-5.4 with reasoning disabled for fastest responses and lowest cost.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (Low Reasoning)
gpt-5.4-low
GPT-5.4 with low reasoning effort for lightweight thinking.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (Medium Reasoning)
gpt-5.4-medium
GPT-5.4 with medium reasoning effort for balanced performance.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (High Reasoning)
gpt-5.4-high
GPT-5.4 with high reasoning effort for complex tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (XHigh Reasoning)
gpt-5.4-xhigh
GPT-5.4 with xhigh reasoning effort for the hardest tasks.
OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 mini
gpt-5.4-mini
GPT-5.4 mini is a faster, more cost-efficient version of GPT-5.4 for well-defined tasks and precise prompts.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (No Reasoning)
gpt-5.4-mini-none
GPT-5.4 mini with reasoning disabled for fastest responses and lowest cost.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (Low Reasoning)
gpt-5.4-mini-low
GPT-5.4 mini with low reasoning effort for lightweight thinking.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (Medium Reasoning)
gpt-5.4-mini-medium
GPT-5.4 mini with medium reasoning effort for balanced performance.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (High Reasoning)
gpt-5.4-mini-high
GPT-5.4 mini with high reasoning effort for complex tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (XHigh Reasoning)
gpt-5.4-mini-xhigh
GPT-5.4 mini with xhigh reasoning effort for the hardest tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 nano
gpt-5.4-nano
GPT-5.4 nano is OpenAI's fastest, cheapest GPT-5.4 model for summarization and classification tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (No Reasoning)
gpt-5.4-nano-none
GPT-5.4 nano with reasoning disabled for fastest responses and lowest cost.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (Low Reasoning)
gpt-5.4-nano-low
GPT-5.4 nano with low reasoning effort for lightweight thinking.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (Medium Reasoning)
gpt-5.4-nano-medium
GPT-5.4 nano with medium reasoning effort for balanced performance.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (High Reasoning)
gpt-5.4-nano-high
GPT-5.4 nano with high reasoning effort for complex tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (XHigh Reasoning)
gpt-5.4-nano-xhigh
GPT-5.4 nano with xhigh reasoning effort for the hardest tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.3 Codex (Low Reasoning)
gpt-5.3-codex-low
GPT-5.3 Codex with low reasoning effort for faster coding tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.3 Codex (Medium Reasoning)
gpt-5.3-codex-medium
GPT-5.3 Codex with medium reasoning effort for balanced performance.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.3 Codex (High Reasoning)
gpt-5.3-codex-high
GPT-5.3 Codex with high reasoning effort for complex coding tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.3 Codex (XHigh Reasoning)
gpt-5.3-codex-xhigh
GPT-5.3 Codex with xhigh reasoning effort for the hardest coding and planning tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 (No Reasoning)
gpt-5.2-none
GPT-5.2 with reasoning disabled for fastest responses and lowest cost.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 (Low Reasoning)
gpt-5.2-low
GPT-5.2 with low reasoning effort for lightweight thinking.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 (Medium Reasoning)
gpt-5.2-medium
GPT-5.2 with medium reasoning effort for balanced performance.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 (High Reasoning)
gpt-5.2-high
GPT-5.2 with high reasoning effort for complex tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.1 (No Reasoning)
gpt-5.1-none
GPT-5.1 with reasoning disabled for fastest responses and lowest cost.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5.1 (Low Reasoning)
gpt-5.1-low
GPT-5.1 with low reasoning effort for lightweight thinking.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5.1 (Medium Reasoning)
gpt-5.1-medium
GPT-5.1 with medium reasoning effort for balanced performance.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5.1 (High Reasoning)
gpt-5.1-high
GPT-5.1 with high reasoning effort for complex tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5
gpt-5
GPT-5 is OpenAI's flagship model for coding, reasoning, and agentic tasks across domains.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5 (High Reasoning)
gpt-5-high
GPT-5 is OpenAI's flagship model for coding, reasoning, and agentic tasks across domains.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5 (Medium Reasoning)
gpt-5-medium
GPT-5 is OpenAI's flagship model for coding, reasoning, and agentic tasks across domains.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5 (Low Reasoning)
gpt-5-low
GPT-5 is OpenAI's flagship model for coding, reasoning, and agentic tasks across domains.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5 (Minimal Reasoning)
gpt-5-minimal
GPT-5 is OpenAI's flagship model for coding, reasoning, and agentic tasks across domains.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5 mini
gpt-5-mini
GPT-5 mini is a faster, more cost-efficient version of GPT-5. It's great for well-defined tasks and precise prompts.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$2.0000
LanguageGPT-5 nano
gpt-5-nano
GPT-5 Nano is OpenAI's fastest, cheapest version of GPT-5. It's great for summarization and classification tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.0500
Output / 1M$0.4000
Language4.1
gpt-4.1
OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.
OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$8.0000
Language4.1 mini
gpt-4.1-mini
GPT 4.1 mini provides a balance between intelligence, speed, and cost that makes it an attractive model for many use cases.
OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.4000
Output / 1M$1.6000
Language4.1 nano
gpt-4.1-nano
GPT-4.1 nano is the fastest, most cost-effective GPT 4.1 model.
OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.4000
Language4o
gpt-4o
Advanced, multimodal flagship model that's cheaper and faster than GPT-4 Turbo
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$10.0000
Language4o-mini
gpt-4o-mini
Affordable and intelligent small model for fast, lightweight tasks. GPT-4o mini is cheaper and more capable than GPT-3.5 Turbo. Currently points to gpt-4o-mini-2024-07-18.
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1500
Output / 1M$0.6000
Languageo3
o3
o3 is a powerful reasoning model designed for complex problem-solving across domains. It combines advanced reasoning capabilities with high performance for demanding tasks.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$8.0000
Languageo3-mini (High Reasoning)
o3-mini-high
Thorough o3-mini model with high reasoning effort. Best for complex tasks requiring deep analysis.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo3-mini (Medium Reasoning)
o3-mini-medium
Balanced o3-mini model with medium reasoning effort. Good for general-purpose tasks requiring moderate analysis.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo3-mini (Low Reasoning)
o3-mini-low
Fast and efficient o3-mini model with low reasoning effort. Optimized for quick responses with basic reasoning.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo4-mini
o4-mini
o4-mini is a compact and efficient model that delivers strong performance for a wide range of tasks. It offers a good balance of capabilities and resource efficiency.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo4-mini (High Reasoning)
o4-mini-high
Thorough o4-mini model with high reasoning effort. Best for complex tasks requiring deep analysis.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo4-mini (Medium Reasoning)
o4-mini-medium
Balanced o4-mini model with medium reasoning effort. Good for general-purpose tasks requiring moderate analysis.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo4-mini (Low Reasoning)
o4-mini-low
Fast and efficient o4-mini model with low reasoning effort. Optimized for quick responses with basic reasoning.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Languageo1
o1
o1 is a reasoning model designed to solve hard problems across domains. The o1 series of models are trained with reinforcement learning to perform complex reasoning. o1 models think before they answer, producing a long internal chain of thought before responding to the user.
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$60.0000
Languageo1-mini
o1-mini
o1-mini is a fast and affordable reasoning model for specialized tasks. The o1-mini series of models are trained with reinforcement learning to perform complex reasoning. o1-mini models think before they answer, producing a long internal chain of thought before responding to the user.
OpenAI
128,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Language4.5
gpt-4.5-preview
This is a research preview of GPT-4.5, OpenAI's largest and most capable GPT model yet. Its deep world knowledge and better understanding of user intent makes it good at creative tasks and agentic planning.
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$75.0000
Output / 1M$150.0000
LanguageGPT-5 2025-08-07
gpt-5-2025-08-07
GPT-5 is OpenAI's flagship model for coding, reasoning, and agentic tasks across domains.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5 mini 2025-08-07
gpt-5-mini-2025-08-07
GPT-5 mini is a faster, more cost-efficient version of GPT-5. It's great for well-defined tasks and precise prompts.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$2.0000
LanguageGPT-5 nano 2025-08-07
gpt-5-nano-2025-08-07
GPT-5 Nano is OpenAI's fastest, cheapest version of GPT-5. It's great for summarization and classification tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.0500
Output / 1M$0.4000
LanguageGPT-5.4 mini 2026-03-17
gpt-5.4-mini-2026-03-17
GPT-5.4 mini is a faster, more cost-efficient version of GPT-5.4 for well-defined tasks and precise prompts.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 nano 2026-03-17
gpt-5.4-nano-2026-03-17
GPT-5.4 nano is OpenAI's fastest, cheapest GPT-5.4 model for summarization and classification tasks.
OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
Language4.1 2025-04-14
gpt-4.1-2025-04-14
OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.
OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$8.0000
Language4.1 mini 2025-04-14
gpt-4.1-mini-2025-04-14
GPT 4.1 mini provides a balance between intelligence, speed, and cost that makes it an attractive model for many use cases.
OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.4000
Output / 1M$1.6000
Language4.1 nano 2025-04-14
gpt-4.1-nano-2025-04-14
GPT-4.1 nano is the fastest, most cost-effective GPT 4.1 model.
OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.4000
Language4o 2024-08-06
gpt-4o-2024-08-06
2024-08-06 version of gpt-4o
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$10.0000
Language4o-mini 2024-07-18
gpt-4o-mini-2024-07-18
2024-07-18 version of gpt-4o-mini
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1500
Output / 1M$0.6000
Languageo1 2024-12-17
o1-2024-12-17
2024-12-17 version of o1
OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$60.0000
Languageo1-mini 2024-09-12
o1-mini-2024-09-12
2024-09-12 version of o1-mini
OpenAI
128,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.1000
Output / 1M$4.4000
Language4 Turbo
gpt-4-turbo
The latest GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more.
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$10.0000
Output / 1M$30.0000
Language4 Turbo Preview
gpt-4-turbo-preview
The latest GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Returns a maximum of 4,096 output tokens. This preview model is not yet suited for production traffic.
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$10.0000
Output / 1M$30.0000
Language4 Vision
gpt-4-vision-preview
GPT-4 with the ability to understand images, in addition to all other GPT-4 Turbo capabilities.
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$10.0000
Output / 1M$30.0000
Language4
gpt-4
More capable than any GPT-3.5 model, able to do more complex tasks, and optimized for chat. Will be updated with our latest model iteration.
OpenAI
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$30.0000
Output / 1M$60.0000
Language4 32K
gpt-4-32k
Same capabilities as the base gpt-4 mode but with 4x the context length. Will be updated with our latest model iteration.
OpenAI
32,768tokens
–Images✓JSON schema✓Function calling
Input / 1M$60.0000
Output / 1M$120.0000
Language4 Turbo 2024-04-09
gpt-4-turbo-2024-04-09
Advanced, multimodal flagship model that's cheaper and faster than GPT-4 Turbo
OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Not listed
Language3.5 Turbo
gpt-3.5-turbo
Most capable GPT-3.5 model and optimized for chat at 1/10th the cost of text-davinci-003. Will be updated with our latest model iteration.
OpenAI
4,096tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$1.5000
Language3.5 Turbo 16K
gpt-3.5-turbo-16k
Same capabilities as the base gpt-3.5-turbo model but with 4x the context length. Will be updated with our latest model iteration.
OpenAI
16,384tokens
–Images✓JSON schema✓Function calling
Not listed
Language4 0613
gpt-4-0613
More capable than any GPT-3.5 model, able to do more complex tasks, and optimized for chat. Will be updated with our latest model iteration.
OpenAI
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$30.0000
Output / 1M$60.0000
Language3.5 Turbo 0613
gpt-3.5-turbo-0613
Most capable GPT-3.5 model and optimized for chat at 1/10th the cost of text-davinci-003. Will be updated with our latest model iteration.
OpenAI
4,096tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.5000
Output / 1M$2.0000
LanguageGPT-5.6 Sol (No Reasoning)
gpt-5.6-sol-none
GPT-5.6 Sol with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (Low Reasoning)
gpt-5.6-sol-low
GPT-5.6 Sol with low reasoning effort for lightweight thinking.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (Medium Reasoning)
gpt-5.6-sol-medium
GPT-5.6 Sol with medium reasoning effort for balanced performance.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (High Reasoning)
gpt-5.6-sol-high
GPT-5.6 Sol with high reasoning effort for complex tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (XHigh Reasoning)
gpt-5.6-sol-xhigh
GPT-5.6 Sol with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Sol (Max Reasoning)
gpt-5.6-sol-max
GPT-5.6 Sol with max reasoning effort for the most demanding tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (No Reasoning)
gpt-5.6-terra-none
GPT-5.6 Terra with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (Low Reasoning)
gpt-5.6-terra-low
GPT-5.6 Terra with low reasoning effort for lightweight thinking.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (Medium Reasoning)
gpt-5.6-terra-medium
GPT-5.6 Terra with medium reasoning effort for balanced performance.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (High Reasoning)
gpt-5.6-terra-high
GPT-5.6 Terra with high reasoning effort for complex tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (XHigh Reasoning)
gpt-5.6-terra-xhigh
GPT-5.6 Terra with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Terra (Max Reasoning)
gpt-5.6-terra-max
GPT-5.6 Terra with max reasoning effort for the most demanding tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (No Reasoning)
gpt-5.6-luna-none
GPT-5.6 Luna with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (Low Reasoning)
gpt-5.6-luna-low
GPT-5.6 Luna with low reasoning effort for lightweight thinking.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (Medium Reasoning)
gpt-5.6-luna-medium
GPT-5.6 Luna with medium reasoning effort for balanced performance.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (High Reasoning)
gpt-5.6-luna-high
GPT-5.6 Luna with high reasoning effort for complex tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (XHigh Reasoning)
gpt-5.6-luna-xhigh
GPT-5.6 Luna with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.6 Luna (Max Reasoning)
gpt-5.6-luna-max
GPT-5.6 Luna with max reasoning effort for the most demanding tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$6.0000
Long context from 272,001 tokens
LanguageGPT-5.5 (No Reasoning)
gpt-5.5-none
GPT-5.5 with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (Low Reasoning)
gpt-5.5-low
GPT-5.5 with low reasoning effort for lightweight thinking.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (Medium Reasoning)
gpt-5.5-medium
GPT-5.5 with medium reasoning effort for balanced performance.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (High Reasoning)
gpt-5.5-high
GPT-5.5 with high reasoning effort for complex tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (XHigh Reasoning)
gpt-5.5-xhigh
GPT-5.5 with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.4 (No Reasoning)
gpt-5.4-none
GPT-5.4 with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (Low Reasoning)
gpt-5.4-low
GPT-5.4 with low reasoning effort for lightweight thinking.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (Medium Reasoning)
gpt-5.4-medium
GPT-5.4 with medium reasoning effort for balanced performance.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (High Reasoning)
gpt-5.4-high
GPT-5.4 with high reasoning effort for complex tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (XHigh Reasoning)
gpt-5.4-xhigh
GPT-5.4 with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 mini
gpt-5.4-mini
GPT-5.4 mini is a faster, more cost-efficient version of GPT-5.4 for well-defined tasks and precise prompts.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (No Reasoning)
gpt-5.4-mini-none
GPT-5.4 mini with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (Low Reasoning)
gpt-5.4-mini-low
GPT-5.4 mini with low reasoning effort for lightweight thinking.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (Medium Reasoning)
gpt-5.4-mini-medium
GPT-5.4 mini with medium reasoning effort for balanced performance.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (High Reasoning)
gpt-5.4-mini-high
GPT-5.4 mini with high reasoning effort for complex tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (XHigh Reasoning)
gpt-5.4-mini-xhigh
GPT-5.4 mini with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 nano
gpt-5.4-nano
GPT-5.4 nano is OpenAI's fastest, cheapest GPT-5.4 model for summarization and classification tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (No Reasoning)
gpt-5.4-nano-none
GPT-5.4 nano with reasoning disabled for fastest responses and lowest cost.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (Low Reasoning)
gpt-5.4-nano-low
GPT-5.4 nano with low reasoning effort for lightweight thinking.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (Medium Reasoning)
gpt-5.4-nano-medium
GPT-5.4 nano with medium reasoning effort for balanced performance.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (High Reasoning)
gpt-5.4-nano-high
GPT-5.4 nano with high reasoning effort for complex tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (XHigh Reasoning)
gpt-5.4-nano-xhigh
GPT-5.4 nano with xhigh reasoning effort for the hardest tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.3 Codex (Low Reasoning)
gpt-5.3-codex-low
GPT-5.3 Codex with low reasoning effort for faster coding tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.3 Codex (Medium Reasoning)
gpt-5.3-codex-medium
GPT-5.3 Codex with medium reasoning effort for balanced performance.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.3 Codex (High Reasoning)
gpt-5.3-codex-high
GPT-5.3 Codex with high reasoning effort for complex coding tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.3 Codex (XHigh Reasoning)
gpt-5.3-codex-xhigh
GPT-5.3 Codex with xhigh reasoning effort for the hardest coding and planning tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 Codex (Low Reasoning)
gpt-5.2-codex-low
GPT-5.2 Codex with low reasoning effort for faster coding tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 Codex (Medium Reasoning)
gpt-5.2-codex-medium
GPT-5.2 Codex with medium reasoning effort for balanced coding performance.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 Codex (High Reasoning)
gpt-5.2-codex-high
GPT-5.2 Codex with high reasoning effort for more demanding coding work.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.2 Codex (XHigh Reasoning)
gpt-5.2-codex-xhigh
GPT-5.2 Codex with extra-high reasoning effort for the hardest coding tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.7500
Output / 1M$14.0000
LanguageGPT-5.1 Codex Max
gpt-5.1-codex-max
GPT-5.1 Codex Max is optimized for long-running agentic coding tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5.1 Codex
gpt-5.1-codex
GPT-5.1 Codex is a coding-optimized GPT-5.1 variant for agentic coding workflows.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-5.1 Codex Mini
gpt-5.1-codex-mini
GPT-5.1 Codex Mini is a smaller, faster Codex model for lighter coding tasks.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$2.0000
LanguageGPT-5 Codex
gpt-5-codex
GPT-5 Codex is a GPT-5 coding model for Codex-oriented development workflows.
OpenAI Codex
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
LanguageGPT-OSS 20B
openai/gpt-oss-20b
OpenAI's flagship open source model, built on a Mixture-of-Experts (MoE) architecture with 20 billion parameters and 32 experts. Features tool use, browser search, code execution, JSON object mode, and reasoning capabilities.
Groq
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.5000
LanguageGPT-OSS 120B
openai/gpt-oss-120b
OpenAI's flagship open source model, built on a Mixture-of-Experts (MoE) architecture with 20 billion parameters and 128 experts. Features tool use, browser search, code execution, JSON object mode, and reasoning capabilities.
Groq
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1500
Output / 1M$0.7500
LanguageKimi K2 Instruct
moonshotai/kimi-k2-instruct
Moonshot AI's state-of-the-art Mixture-of-Experts (MoE) language model with 1 trillion total parameters and 32 billion activated parameters. Designed for agentic intelligence, it excels at tool use, coding, and autonomous problem-solving across diverse domains.
Groq
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$3.0000
LanguageLlama 4 Maverick
meta-llama/llama-4-maverick-17b-128e-instruct
Llama 4 Maverick
Groq
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$0.7700
LanguageLlama 4 Scout
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout
Groq
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1100
Output / 1M$0.3400
LanguageDeepSeek R1 Distilled Llama 70B
deepseek-r1-distill-llama-70b
DeepSeek R1 Distilled Llama 70B
Groq
128,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$8.0000
Output / 1M$8.0000
LanguageDeepSeek R1 Distilled Llama 70B SpecDec
deepseek-r1-distill-llama-70b-specdec
DeepSeek R1 Distilled Llama 70B SpecDec
Groq
128,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$8.0000
Output / 1M$8.0000
LanguageLlama 3.1 405B Reasoning
llama-3.1-405b-reasoning
Llama 3.1 405B Reasoning
Groq
131,072tokens
–Images✓JSON schema–Function calling
Input / 1M$0.5900
Output / 1M$0.7900
LanguageLlama 3.3 70B Versatile
llama-3.3-70b-versatile
Llama 3.3 70B Versatile
Groq
32,768tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.5900
Output / 1M$0.7900
LanguageLlama 3.3 70B SpecDec
llama-3.3-70b-specdec
Llama 3.3 70B SpecDec
Groq
8,192tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageLlama 3.1 70B Versatile (Tool Use Preview)
llama3-groq-70b-8192-tool-use-preview
Llama 3.1 70B Versatile (Tool Use Preview)
Groq
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.5900
Output / 1M$0.7900
LanguageLlama 3.1 70B Versatile
llama-3.1-70b-versatile
Llama 3.1 70B Versatile
Groq
131,072tokens
–Images✓JSON schema–Function calling
Input / 1M$0.5900
Output / 1M$0.7900
LanguageLlama 3.1 8B Instant (Tool Use Preview)
llama3-groq-8b-8192-tool-use-preview
Llama 3.1 8B Instant (Tool Use Preview)
Groq
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.0500
Output / 1M$0.1000
LanguageLlama 3.1 8B Instant
llama-3.1-8b-instant
Llama 3.1 8B Instant
Groq
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.0500
Output / 1M$0.1000
LanguageLLaMA3-70b
llama3-70b-8192
LLaMA3-70b
Groq
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.5900
Output / 1M$0.7900
LanguageLLaMA3-8b
llama3-8b-8192
LLaMA3-8b
Groq
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.0500
Output / 1M$0.1000
LanguageLLaMA2-70b
llama2-70b-4096
LLaMA2-70b
Groq
4,096tokens
–Images–JSON schema–Function calling
Input / 1M$0.6400
Output / 1M$0.8000
LanguageMixtral-8x7b
mixtral-8x7b-32768
Mixtral-8x7b
Groq
32,768tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.2700
Output / 1M$0.2700
LanguageGemma-7b-it
gemma-7b-it
Gemma-7b-it
Groq
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.1000
LanguageGemini 3.5 Flash (Medium Thinking)
gemini-3.5-flash
Gemini 3.5 Flash with Medium Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.5000
Output / 1M$9.0000
LanguageGemini 3.5 Flash (High Thinking)
gemini-3.5-flash-high
Gemini 3.5 Flash with High Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.5000
Output / 1M$9.0000
LanguageGemini 3.5 Flash (Low Thinking)
gemini-3.5-flash-low
Gemini 3.5 Flash with Low Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.5000
Output / 1M$9.0000
LanguageGemini 3.5 Flash (Minimal Thinking)
gemini-3.5-flash-minimal
Gemini 3.5 Flash with Minimal Thinking enabled for fastest responses.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.5000
Output / 1M$9.0000
LanguageGemini 3.1 Pro (High Thinking)
gemini-3.1-pro-preview
Gemini 3.1 Pro with High Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$12.0000
Long context from 200,000 tokens
LanguageGemini 3.1 Pro (Low Thinking)
gemini-3.1-pro-preview-low
Gemini 3.1 Pro with Low Thinking enabled for faster responses.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$12.0000
Long context from 200,000 tokens
LanguageGemini 3.1 Pro (Custom Tools Preview)
gemini-3.1-pro-preview-customtools
Gemini 3.1 Pro preview variant optimized for custom tool use in agentic workflows.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$12.0000
Long context from 200,000 tokens
LanguageGemini 3.1 Flash-Lite (High Thinking)
gemini-3.1-flash-lite
Gemini 3.1 Flash-Lite with High Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$1.5000
LanguageGemini 3.1 Flash-Lite (Medium Thinking)
gemini-3.1-flash-lite-medium
Gemini 3.1 Flash-Lite with Medium Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$1.5000
LanguageGemini 3.1 Flash-Lite (Low Thinking)
gemini-3.1-flash-lite-low
Gemini 3.1 Flash-Lite with Low Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$1.5000
LanguageGemini 3.1 Flash-Lite (Minimal Thinking)
gemini-3.1-flash-lite-minimal
Gemini 3.1 Flash-Lite with Minimal Thinking enabled for fastest responses.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$1.5000
LanguageGemini 3 Pro (High Thinking)
gemini-3-pro-preview
Gemini 3 Pro with High Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$12.0000
Long context from 200,000 tokens
LanguageGemini 3 Pro (Low Thinking)
gemini-3-pro-preview-low
Gemini 3 Pro with Low Thinking enabled for faster responses.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$12.0000
Long context from 200,000 tokens
LanguageGemini 3 Flash (High Thinking)
gemini-3-flash-preview
Gemini 3 Flash with High Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$3.0000
LanguageGemini 3 Flash (Medium Thinking)
gemini-3-flash-preview-medium
Gemini 3 Flash with Medium Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$3.0000
LanguageGemini 3 Flash (Low Thinking)
gemini-3-flash-preview-low
Gemini 3 Flash with Low Thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$3.0000
LanguageGemini 3 Flash (Minimal Thinking)
gemini-3-flash-preview-minimal
Gemini 3 Flash with Minimal Thinking enabled for fastest responses.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$3.0000
LanguageGemini 2.5 Pro
gemini-2.5-pro
Gemini 2.5 Pro
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
Long context from 200,000 tokens
LanguageGemini 2.5 Pro (Thinking)
gemini-2.5-pro-thinking
Gemini 2.5 Pro with dynamic thinking output enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$10.0000
Long context from 200,000 tokens
LanguageGemini 2.5 Computer Use Preview (10-2025)
gemini-2.5-computer-use-preview-10-2025
Gemini 2.5 Computer Use model optimized for browser automation tasks.
Gemini 2.5 Flash-Lite Preview 09-2025 with dynamic thinking enabled.
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.4000
LanguageGemini 2.0 Flash
gemini-2.0-flash-001
Gemini 2.0 Flash
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.3750
Output / 1M$1.5000
LanguageGemini 2.0 Flash Experimental
gemini-2.0-flash-exp
Gemini 2.0 Flash Experimental
Google
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.0000
Output / 1M$0.0000
LanguageGemini 1.0 Pro
gemini-pro
Gemini 1.0 Pro
Google
32,000tokens
–Images–JSON schema–Function calling
Input / 1M$1.2500
Output / 1M$3.7500
LanguageClaude Fable 5
claude-fable-5
Anthropic's most capable widely released model for demanding reasoning and long-horizon agentic work, with a 1M-token context window.
Anthropic Claude
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$10.0000
Output / 1M$50.0000
LanguageClaude Opus 4.8
claude-opus-4-8
Anthropic's frontier Opus model for coding, agentic workflows, and high-stakes enterprise tasks with adaptive thinking and a 1M-token context window.
Anthropic Claude
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
LanguageClaude Opus 4.7
claude-opus-4-7
Anthropic's latest Opus model for advanced coding and long-running agentic workflows. Announced April 16, 2026 with the same base pricing as Opus 4.6.
Anthropic Claude
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
Long context from 200,000 tokens
LanguageClaude Opus 4.6
claude-opus-4-6
Anthropic's most capable Claude model, tuned for stronger coding and agentic reliability with reduced reward-hacking behavior on long-running tasks.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
Long context from 200,000 tokens
LanguageClaude Opus 4.5
claude-opus-4-5-20251101
Anthropic's most intelligent and capable model. State-of-the-art for coding, agents, and computer use with industry-leading performance on complex reasoning tasks.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
LanguageClaude Sonnet 4.6
claude-sonnet-4-6
Anthropic's most intelligent Sonnet model with superior coding and reasoning performance, agentic reliability improvements, and 200K context support.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 4.5 Sonnet
claude-sonnet-4-5-20250929
Anthropic's hybrid-reasoning model. Seamlessly switches between rapid standard responses and extended thinking mode for visible step-by-step reasoning. Features a 200,000-token context window (expandable to 1M) with state-of-the-art coding performance and multimodal capabilities.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 4.5 Haiku
claude-haiku-4-5
Anthropic's fastest Claude 4.5 model optimized for rapid responses while retaining multimodal support and extended context.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$5.0000
LanguageClaude Opus 4.1
claude-opus-4-1-20250805
Anthropic's most capable and intelligent model yet. Claude Opus 4.1 sets new standards in complex reasoning and advanced coding.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$75.0000
LanguageClaude Opus 4
claude-opus-4-20250514
Anthropic's most capable model with highest level of intelligence and capability. Features extended thinking and priority tier access.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$75.0000
LanguageClaude Sonnet 4
claude-sonnet-4-20250514
Anthropic's high-performance model with balanced intelligence and speed. Features extended thinking and priority tier access.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 3.7 Sonnet
claude-3-7-sonnet-20250219
Anthropic's most intelligent model. Highest level of intelligence and capability with toggleable extended thinking. This is the latest version of the model.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageClaude 3.5 Sonnet (V2)
claude-3-5-sonnet-20241022
Anthropic's previous most intelligent model. High level of intelligence and capability.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageClaude 3.5 Sonnet (V1)
claude-3-5-sonnet-20240620
Anthropic's previous most intelligent model. High level of intelligence and capability.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageClaude 3.5 Haiku
claude-3-5-haiku-20241022
Anthropic's fastest model that can execute lightweight actions, with industry-leading speed.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.8000
Output / 1M$4.0000
LanguageClaude 3 Opus
claude-3-opus-20240229
Most powerful model for highly complex tasks, offering top-level performance with multilingual and vision capabilities.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$75.0000
LanguageClaude 3 Sonnet
claude-3-sonnet-20240229
Ideal balance of intelligence and speed for enterprise workloads, with multilingual and vision support.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageClaude 3 Haiku
claude-3-haiku-20240307
Fastest and most compact model for near-instant responsiveness, includes multilingual and vision capabilities.
Anthropic Claude
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$1.2500
LanguageSonar
sonar
Lightweight, cost-effective search model with grounding. Best suited for quick factual queries, topic summaries, product comparisons, and current events.
Perplexity AI
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$1.0000
Output / 1M$1.0000
LanguageSonar Pro
sonar-pro
Advanced search offering with grounding, supporting complex queries and follow-ups. Ideal for detailed information retrieval and synthesis.
Perplexity AI
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageSonar Reasoning
sonar-reasoning
Fast, real-time reasoning model designed for problem-solving with search. Excellent for complex analyses requiring step-by-step thinking.
Perplexity AI
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$1.0000
Output / 1M$5.0000
LanguageSonar Deep Research
sonar-deep-research
Expert-level research model conducting exhaustive searches and generating comprehensive reports. Ideal for in-depth analysis and detailed topic reports.
Perplexity AI
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$2.0000
Output / 1M$8.0000
LanguageLlama-3.1-Sonar-Small (8B)
llama-3.1-sonar-small-128k-online
Meta's Llama-3.1-Sonar-Small model with 8 billion parameters for chat use cases.
Perplexity AI
127,072tokens
–Images–JSON schema–Function calling
Input / 1M$0.2000
Output / 1M$0.2000
LanguageLlama-3.1-Sonar-Large (70B)
llama-3.1-sonar-large-128k-online
Meta's Llama-3.1-Sonar-Large model with 70 billion parameters for chat use cases.
Perplexity AI
127,072tokens
–Images–JSON schema–Function calling
Input / 1M$1.0000
Output / 1M$1.0000
LanguageLlama-3.1-Sonar-Huge (405B)
llama-3.1-sonar-huge-128k-online
Meta's Llama-3.1-Sonar-Huge model with 405 billion parameters for chat use cases.
Perplexity AI
127,072tokens
–Images–JSON schema–Function calling
Input / 1M$5.0000
Output / 1M$5.0000
LanguageClaude Opus 4.8
anthropic.claude-opus-4-8
Anthropic's Claude Opus 4.8 model on Amazon Bedrock
Amazon Bedrock
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
LanguageClaude Opus 4.6
anthropic.claude-opus-4-6-v1
Anthropic's Claude Opus 4.6 model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Not listed
LanguageClaude Opus 4.5
anthropic.claude-opus-4-5-20251101-v1:0
Anthropic's Claude Opus 4.5 model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Not listed
LanguageClaude Opus 4.1
anthropic.claude-opus-4-1-20250805-v1:0
Anthropic's Claude Opus 4.1 model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$75.0000
LanguageClaude Opus 4
anthropic.claude-opus-4-20250514-v1:0
Anthropic's Claude Opus 4 model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$15.0000
Output / 1M$75.0000
LanguageClaude Sonnet 4.6
anthropic.claude-sonnet-4-6
Anthropic's Claude Sonnet 4.6 model on Amazon Bedrock
Amazon Bedrock
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude Sonnet 4.5
anthropic.claude-sonnet-4-5-20250929-v1:0
Anthropic's Claude 4.5 Sonnet model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude Sonnet 4
anthropic.claude-sonnet-4-20250514-v1:0
Anthropic's Claude Sonnet 4 model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude Haiku 4.5
anthropic.claude-haiku-4-5-20251001-v1:0
Anthropic's Claude 4.5 Haiku model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Not listed
LanguageClaude 3.7 Sonnet
anthropic.claude-3-7-sonnet-20250219-v1:0
Anthropic's Claude 3.7 Sonnet model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 3.5 Sonnet (V2)
anthropic.claude-3-5-sonnet-20241022-v2:0
Anthropic's Claude 3.5 Sonnet model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 3.5 Sonnet
anthropic.claude-3-5-sonnet-20240620-v1:0
Anthropic's Claude 3.5 Sonnet model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 3 Sonnet
anthropic.claude-3-sonnet-20240229-v1:0
Anthropic's Claude 3 Sonnet model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema–Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 200,000 tokens
LanguageClaude 3.5 Haiku
anthropic.claude-3-5-haiku-20241022-v1:0
Anthropic's Claude 3.5 Haiku model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.8000
Output / 1M$4.0000
LanguageClaude 3 Haiku
anthropic.claude-3-haiku-20240307-v1:0
Anthropic's Claude 3 Haiku model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2500
Output / 1M$1.2500
LanguageClaude 3 Opus
anthropic.claude-3-opus-20240229-v1:0
Anthropic's Claude 3 Opus model on Amazon Bedrock
Amazon Bedrock
200,000tokens
✓Images✓JSON schema–Function calling
Input / 1M$15.0000
Output / 1M$75.0000
LanguageKimi K2.5
moonshotai.kimi-k2.5
Moonshot AI's Kimi K2.5 model on Amazon Bedrock
Amazon Bedrock
256,000tokens
✓Images✓JSON schema✓Function calling
Not listed
LanguageKimi K2 Thinking
moonshot.kimi-k2-thinking
Moonshot AI's Kimi K2 Thinking model on Amazon Bedrock
Amazon Bedrock
256,000tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageMiniMax M2.5
minimax.minimax-m2.5
MiniMax M2.5 on Amazon Bedrock
Amazon Bedrock
196,608tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageMiniMax M2.1
minimax.minimax-m2.1
MiniMax M2.1 on Amazon Bedrock
Amazon Bedrock
196,608tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageMiniMax M2
minimax.minimax-m2
MiniMax M2 on Amazon Bedrock
Amazon Bedrock
196,608tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageGPT-OSS 20B
openai.gpt-oss-20b-1:0
OpenAI's GPT-OSS 20B model on Amazon Bedrock for efficient text generation and coding.
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.0700
Output / 1M$0.3000
LanguageGPT-OSS 120B
openai.gpt-oss-120b-1:0
OpenAI's GPT-OSS 120B general-purpose model on Amazon Bedrock for text generation, coding, and reasoning.
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.1500
Output / 1M$0.6000
LanguageGPT-5.5
openai.gpt-5.5
OpenAI's GPT-5.5 frontier model on Amazon Bedrock through the OpenAI-compatible Responses API. Available in us-east-2.
Amazon Bedrock
272,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.4
openai.gpt-5.4
OpenAI's GPT-5.4 frontier model on Amazon Bedrock through the OpenAI-compatible Responses API. Available in us-east-2 and us-west-2.
Amazon Bedrock
272,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
LanguageLlama 3 8B Instruct
meta.llama3-8b-instruct-v1:0
Meta's Llama 3 8B Instruct model on Amazon Bedrock
Amazon Bedrock
4,096tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3 70B Instruct
meta.llama3-70b-instruct-v1:0
Meta's Llama 3 70B Instruct model on Amazon Bedrock
Amazon Bedrock
4,096tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.1 8B Instruct
meta.llama3-1-8b-instruct-v1:0
Meta's Llama 3.1 8B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.1 70B Instruct
meta.llama3-1-70b-instruct-v1:0
Meta's Llama 3.1 70B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.1 405B Instruct
meta.llama3-1-405b-instruct-v1:0
Meta's Llama 3.1 405B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.2 1B Instruct
us.meta.llama3-2-1b-instruct-v1:0
Meta's Llama 3.2 1B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.2 3B Instruct
us.meta.llama3-2-3b-instruct-v1:0
Meta's Llama 3.2 3B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.2 11B Instruct
us.meta.llama3-2-11b-instruct-v1:0
Meta's Llama 3.2 11B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images–JSON schema–Function calling
Not listed
LanguageLlama 3.2 90B Instruct
us.meta.llama3-2-90b-instruct-v1:0
Meta's Llama 3.2 90B Instruct model on Amazon Bedrock
Amazon Bedrock
128,000tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageGPT-4.1
gpt-4.1
Most capable GPT-4.1 model for tasks requiring deep understanding and advanced reasoning.
Azure OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$8.0000
LanguageGPT-4.1 Mini
gpt-4.1-mini
Smaller, faster version of GPT-4.1 optimized for efficiency.
Azure OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.4000
Output / 1M$1.6000
LanguageGPT-4.1 Nano
gpt-4.1-nano
Smallest version of GPT-4.1 optimized for speed and cost efficiency.
Azure OpenAI
1,047,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.4000
LanguageGPT-5.5 (No Reasoning)
gpt-5.5-none
GPT-5.5 with reasoning disabled for fastest responses and lowest cost.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (Low Reasoning)
gpt-5.5-low
GPT-5.5 with low reasoning for lightweight thinking.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (Medium Reasoning)
gpt-5.5-medium
GPT-5.5 with medium reasoning for balanced performance.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (High Reasoning)
gpt-5.5-high
GPT-5.5 with high reasoning for complex tasks.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.5 (XHigh Reasoning)
gpt-5.5-xhigh
GPT-5.5 with xhigh reasoning for the hardest tasks.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$30.0000
LanguageGPT-5.4 (No Reasoning)
gpt-5.4-none
GPT-5.4 with reasoning disabled for fastest responses and lowest cost.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (Low Reasoning)
gpt-5.4-low
GPT-5.4 with low reasoning for lightweight thinking.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (Medium Reasoning)
gpt-5.4-medium
GPT-5.4 with medium reasoning for balanced performance.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (High Reasoning)
gpt-5.4-high
GPT-5.4 with high reasoning for complex tasks.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 (XHigh Reasoning)
gpt-5.4-xhigh
GPT-5.4 with xhigh reasoning for the hardest tasks.
Azure OpenAI
1,050,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.5000
Output / 1M$15.0000
Long context from 272,000 tokens
LanguageGPT-5.4 mini
gpt-5.4-mini
GPT-5.4 mini is a faster, more cost-efficient version of GPT-5.4 for well-defined tasks and precise prompts.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (No Reasoning)
gpt-5.4-mini-none
GPT-5.4 mini with reasoning disabled for fastest responses and lowest cost.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (Low Reasoning)
gpt-5.4-mini-low
GPT-5.4 mini with low reasoning for lightweight thinking.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (Medium Reasoning)
gpt-5.4-mini-medium
GPT-5.4 mini with medium reasoning for balanced performance.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (High Reasoning)
gpt-5.4-mini-high
GPT-5.4 mini with high reasoning for complex tasks.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 mini (XHigh Reasoning)
gpt-5.4-mini-xhigh
GPT-5.4 mini with xhigh reasoning for the hardest tasks.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$4.5000
LanguageGPT-5.4 nano
gpt-5.4-nano
GPT-5.4 nano is OpenAI's fastest, cheapest GPT-5.4 model for summarization and classification tasks.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (No Reasoning)
gpt-5.4-nano-none
GPT-5.4 nano with reasoning disabled for fastest responses and lowest cost.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (Low Reasoning)
gpt-5.4-nano-low
GPT-5.4 nano with low reasoning for lightweight thinking.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (Medium Reasoning)
gpt-5.4-nano-medium
GPT-5.4 nano with medium reasoning for balanced performance.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (High Reasoning)
gpt-5.4-nano-high
GPT-5.4 nano with high reasoning for complex tasks.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-5.4 nano (XHigh Reasoning)
gpt-5.4-nano-xhigh
GPT-5.4 nano with xhigh reasoning for the hardest tasks.
Azure OpenAI
400,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$1.2500
LanguageGPT-4o
gpt-4o
Latest large GA model with structured outputs, text/image processing, enhanced accuracy and superior performance in non-English languages and vision tasks.
Azure OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$15.0000
LanguageGPT-4o mini
gpt-4o-mini
Latest small GA model optimized for fast, inexpensive tasks. Supports text and image processing, JSON Mode, and parallel function calling.
Azure OpenAI
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1500
Output / 1M$0.6000
Languageo1
o1
o1 is a reasoning model designed to solve hard problems across domains. The o1 series of models are trained with reinforcement learning to perform complex reasoning. o1 models think before they answer, producing a long internal chain of thought before responding to the user.
Azure OpenAI
200,000tokens
✓Images✓JSON schema✓Function calling
Not listed
Languageo1-mini
o1-mini
o1-mini is a fast and affordable reasoning model for specialized tasks. The o1-mini series of models are trained with reinforcement learning to perform complex reasoning. o1-mini models think before they answer, producing a long internal chain of thought before responding to the user.
Azure OpenAI
128,000tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageGPT-4
gpt-4
Most capable GPT-4 model for tasks requiring deep understanding and advanced reasoning.
Azure OpenAI
8,192tokens
–Images✓JSON schema✓Function calling
Input / 1M$30.0000
Output / 1M$60.0000
LanguageGPT-3.5 Turbo
gpt-35-turbo
Most capable GPT-3.5 model, optimized for chat at 1/10th the cost of GPT-4.
Azure OpenAI
16,385tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$1.5000
LanguageGrok 4.5 (Low Reasoning)
grok-4.5-low
Grok 4.5 with low reasoning effort. Supports text and image input, 500k token context window, function calling, structured outputs, and configurable reasoning.
xAI
500,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$6.0000
Long context from 200,000 tokens
LanguageGrok 4.5 (Medium Reasoning)
grok-4.5-medium
Grok 4.5 with medium reasoning effort. Supports text and image input, 500k token context window, function calling, structured outputs, and configurable reasoning.
xAI
500,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$6.0000
Long context from 200,000 tokens
LanguageGrok 4.5 (High Reasoning)
grok-4.5-high
Grok 4.5 with high reasoning effort. Supports text and image input, 500k token context window, function calling, structured outputs, and configurable reasoning.
xAI
500,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$6.0000
Long context from 200,000 tokens
LanguageGrok 4.3
grok-4.3
Grok 4.3 reasoning model. Supports text and image input, 1M token context window, function calling, structured outputs, and reasoning.
xAI
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$2.5000
LanguageGrok 4.20 (Reasoning)
grok-4.20-0309-reasoning
Grok 4.20 reasoning model. Supports text and image input, 2M token context window, function calling, structured outputs, and reasoning.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$6.0000
LanguageGrok 4.20 (Non-Reasoning)
grok-4.20-0309-non-reasoning
Grok 4.20 non-reasoning model. Supports text and image input, 2M token context window, function calling, and structured outputs.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$6.0000
LanguageGrok 4.20 Multi-Agent
grok-4.20-multi-agent-0309
Grok 4.20 Multi-Agent model for realtime multi-agent research. Supports structured outputs, reasoning, responses API, and xAI's built-in research tool loop.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$2.0000
Output / 1M$6.0000
LanguageGrok 4.1 Fast
grok-4-1-fast-reasoning
Grok 4.1 Fast Reasoning model. Supports text and image input, 2M token context window, reasoning, function calling, and structured outputs.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$0.5000
Long context from 128,000 tokens
LanguageGrok 4.1 Fast (Non-Reasoning)
grok-4-1-fast-non-reasoning
Grok 4.1 Fast Non Reasoning model. Supports text and image input, 2M token context window, function calling, and structured outputs.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$0.5000
Long context from 128,000 tokens
LanguageGrok 4 Fast
grok-4-fast-reasoning
Grok 4 Fast Reasoning model. Supports text and image input, 2M token context window, reasoning, function calling, and structured outputs.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$0.5000
Long context from 128,000 tokens
LanguageGrok 4 Fast (Non-Reasoning)
grok-4-fast-non-reasoning
Grok 4 Fast Non Reasoning model. Supports text and image input, 2M token context window, function calling, and structured outputs.
xAI
2,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$0.5000
Long context from 128,000 tokens
LanguageGrok 4 (July 2024)
grok-4-0709
Grok 4 model (July 2024). Supports text and image input, 256,000 token context window, advanced reasoning, function calling, and structured outputs.
xAI
256,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
Long context from 128,000 tokens
LanguageGrok 3
grok-3
Grok 3 model with high performance capabilities. Choose this for reduced cost compared to grok-3-fast.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageGrok 3 Latest
grok-3-latest
Latest version of Grok 3 model with high performance capabilities.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$3.0000
Output / 1M$15.0000
LanguageGrok 3 Fast
grok-3-fast
Same as Grok 3 model but optimized for latency-sensitive applications. Choose this for better response time at higher cost.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
LanguageGrok 3 Fast Latest
grok-3-fast-latest
Latest faster version of Grok 3 model with optimized response time.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$25.0000
LanguageGrok 3 Mini
grok-3-mini
Lightweight version of Grok 3 model with lower cost and good performance.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$0.5000
LanguageGrok 3 Mini Latest
grok-3-mini-latest
Latest lightweight version of Grok 3 model with lower cost and good performance.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$0.5000
LanguageGrok 3 Mini Fast
grok-3-mini-fast
Faster lightweight version of Grok 3 model with balanced cost and performance.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.6000
Output / 1M$4.0000
LanguageGrok 3 Mini Fast Latest
grok-3-mini-fast-latest
Latest faster lightweight version of Grok 3 model with balanced cost and performance.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.6000
Output / 1M$4.0000
LanguageGrok Beta
grok-beta
Comparable performance to Grok 2 but with improved efficiency, speed and capabilities.
xAI
131,072tokens
–Images✓JSON schema✓Function calling
Input / 1M$5.0000
Output / 1M$15.0000
LanguageGrok Vision Beta
grok-vision-beta
Comparable performance to Grok 2 but with improved efficiency, speed and capabilities and with ability to process images.
xAI
8,192tokens
✓Images✓JSON schema–Function calling
Input / 1M$5.0000
Output / 1M$15.0000
LanguageMiniMax M3
accounts/fireworks/models/minimax-m3
MiniMax M3 is a native multimodal model with 512K context, MiniMax Sparse Attention, and strong long-horizon agentic coding performance.
Fireworks
512,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$1.2000
LanguageMiniMax M3 (Non-thinking)
accounts/fireworks/models/minimax-m3-non-thinking
MiniMax M3 with reasoning disabled for lower-latency chat and code-completion scenarios.
Fireworks
512,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$1.2000
LanguageMiniMax M2.5
accounts/fireworks/models/minimax-m2p5
MiniMax M2.5 is a 228.7B mixture-of-experts model extensively trained with reinforcement learning for state-of-the-art coding, agentic tool use, search, and office work.
Fireworks
196,608tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$1.2000
LanguageMiniMax M2.7
accounts/fireworks/models/minimax-m2p7
MiniMax M2.7 is MiniMax's latest M2 model available through Fireworks.
Fireworks
196,608tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$1.2000
LanguageDeepSeek V3.2
accounts/fireworks/models/deepseek-v3p2
DeepSeek V3.2 is a model from Deepseek that harmonizes high computational efficiency with superior reasoning and agent performance.
Fireworks
163,800tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.2000
Output / 1M$1.2000
LanguageDeepSeek V4 Pro
accounts/fireworks/models/deepseek-v4-pro
DeepSeek V4 Pro is DeepSeek's flagship open-source MoE model for frontier reasoning, advanced coding, and long-context agentic workflows on Fireworks.
Fireworks
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.7400
Output / 1M$3.4800
LanguageDeepSeek V4 Flash
accounts/fireworks/models/deepseek-v4-flash
DeepSeek V4 Flash is DeepSeek's fast, cost-efficient open-source MoE model for long-context reasoning, coding, and high-volume agentic workflows on Fireworks.
Fireworks
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1400
Output / 1M$0.2800
LanguageGLM-5
accounts/fireworks/models/glm-5
GLM-5 is Z.ai's state-of-the-art model for complex systems engineering, coding, and long-horizon agentic tasks.
Fireworks
202,800tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$3.2000
LanguageGLM 5.1
accounts/fireworks/models/glm-5p1
GLM 5.1 is Z.ai's latest model with enhanced capabilities for complex systems engineering, coding, and long-horizon agentic tasks with an expanded 202k context window.
Fireworks
202,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.4000
Output / 1M$4.4000
LanguageGLM 5.2
accounts/fireworks/models/glm-5p2
GLM 5.2 is Z.ai's flagship model for long-horizon coding and agentic engineering tasks, with a 1M-token context window and multi-effort reasoning.
Fireworks
1,040,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.4000
Output / 1M$4.4000
LanguageKimi K2 Thinking
accounts/fireworks/models/kimi-k2-thinking
Kimi K2 Thinking is the latest version of Moonshot AI's open-source thinking model, designed for advanced reasoning tasks. It interleaves step-by-step chain-of-thought reasoning with autonomous tool use, achieving strong performance across benchmarks like HLE, AIME25, and BrowseComp.
Fireworks
256,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.6000
Output / 1M$2.5000
LanguageKimi K2.5
accounts/fireworks/models/kimi-k2p5
Kimi K2.5 is Moonshot AI's flagship agentic model, unifying vision and text capabilities along with both thinking and non-thinking execution modes within a single framework.
Fireworks
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.6000
Output / 1M$3.0000
LanguageKimi K2.6
accounts/fireworks/models/kimi-k2p6
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
Fireworks
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.9500
Output / 1M$4.0000
LanguageKimi K2.7 Code
accounts/fireworks/models/kimi-k2p7-code
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6, with improvements for long-horizon software engineering workflows and token efficiency.
Fireworks
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.9500
Output / 1M$4.0000
LanguageGPT-OSS 20B
accounts/fireworks/models/gpt-oss-20b
A compact, open-weight language model optimized for low-latency and resource-constrained environments, including local and edge deployments. It shares the same Harmony training foundation and capabilities as 120B, with faster inference and easier deployment that is ideal for specialized or offline use cases, fast responsive performance, chain-of-thought output and adjustable reasoning levels, and agentic workflows.
Fireworks
128,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.0700
Output / 1M$0.3000
LanguageGPT-OSS 120B
accounts/fireworks/models/gpt-oss-120b
A high-performance, open-weight language model designed for production-grade, general-purpose use cases. It fits on a single H100 GPU, making it accessible without requiring multi-GPU infrastructure. Trained on the Harmony response format, it excels at complex reasoning and supports configurable reasoning effort, full chain-of-thought transparency for easier debugging and trust, and native agentic capabilities for function calling, tool use, and structured outputs.
The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding.
The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding.
Fireworks
128,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1500
Output / 1M$0.6000
LanguageQwen3 235B A22B
accounts/fireworks/models/qwen3-235b-a22b
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models
Fireworks
32,768tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.1000
LanguageQwen3.6 Plus
accounts/fireworks/models/qwen3p6-plus
Qwen3.6 Plus is the latest generation of large language models in the Qwen series, offering advanced reasoning and vision capabilities with serverless deployment.
Fireworks
131,072tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$3.0000
LanguageQwen3.7 Plus
accounts/fireworks/models/qwen3p7-plus
Qwen3.7 Plus is Alibaba's latest flagship multimodal model available through Fireworks serverless inference.
Fireworks
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.4000
Output / 1M$1.6000
LanguageDeepSeek R1
accounts/fireworks/models/deepseek-r1
DeepSeek R1 is a large language model optimized for instruction following and coding tasks.
Fireworks
160,000tokens
–Images–JSON schema–Function calling
Input / 1M$3.0000
Output / 1M$8.0000
LanguageDeepSeek V3 03-24
accounts/fireworks/models/deepseek-v3-0324
DeepSeek V3 is a large language model optimized for instruction following. This model is the version of the DeepSeek V3 model as of 3/24/2025.
Fireworks
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$1.2000
Output / 1M$1.2000
LanguageDeepSeek V3
accounts/fireworks/models/deepseek-v3
DeepSeek V3 is a large language model optimized for instruction following.
Fireworks
128,000tokens
–Images–JSON schema–Function calling
Input / 1M$0.9000
Output / 1M$0.9000
LanguageLlama 3.3 70B Instruct
accounts/fireworks/models/llama-v3p3-70b-instruct
Llama 3.3 70B Instruct is a large language model that is optimized for instruction following.
Llama 3.1 405B Instruct is a large language model that is optimized for instruction following.
Fireworks
128,000tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageLlama 3.1 70B Instruct
accounts/fireworks/models/llama-v3p1-70b-instruct
Llama 3.1 70B Instruct is a large language model that is optimized for instruction following.
Fireworks
128,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.9000
Output / 1M$0.9000
LanguageKimi K2.6
moonshotai/Kimi-K2.6
Moonshot AI's latest Kimi K2 model available through Together AI.
Together AI
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2000
Output / 1M$4.5000
LanguageKimi K2 Thinking
moonshotai/Kimi-K2-Thinking
Moonshot AI's Kimi K2 thinking model available through Together AI.
Together AI
262,144tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageMiniMax M2.5
MiniMaxAI/MiniMax-M2.5
MiniMax M2.5 available through Together AI.
Together AI
196,608tokens
–Images✓JSON schema✓Function calling
Not listed
LanguageMiniMax M2.7
MiniMaxAI/MiniMax-M2.7
MiniMax's latest M2 model available through Together AI.
Together AI
202,752tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$1.2000
LanguageMiniMax M3
MiniMaxAI/MiniMax-M3
MiniMax M3 is MiniMax's frontier open-weight model available through Together AI, combining coding, agentic capability, 1M context, and native multimodality.
Together AI
1,048,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.3000
Output / 1M$1.2000
LanguageDeepSeek V4-Pro
deepseek-ai/DeepSeek-V4-Pro
DeepSeek's latest V4-Pro model available through Together AI.
Together AI
512,000tokens
–Images✓JSON schema✓Function calling
Input / 1M$2.1000
Output / 1M$4.4000
LanguageGLM-5.2
zai-org/GLM-5.2
Z.ai's GLM-5.2 model available through Together AI for long-horizon coding and agentic engineering tasks.
Together AI
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.4000
Output / 1M$4.4000
LanguageGLM-5.2
zai-org/GLM-5.2
GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks, with a solid 1M-token context window.
DeepInfra
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.3950
Output / 1M$4.5000
LanguageNVIDIA-Nemotron-3-Ultra-550B-A55B
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows.
DeepInfra
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.5000
Output / 1M$2.5000
LanguageNemotron-3-Nano-Omni-30B-A3B-Reasoning
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning
Nemotron 3 Nano Omni is an open multimodal model built on a hybrid Mixture-of-Experts architecture for image, video, audio, and text inputs.
DeepInfra
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.2000
Output / 1M$0.8000
LanguageDeepSeek V4 Flash
deepseek-ai/DeepSeek-V4-Flash
DeepSeek V4 Flash is an efficiency-focused MoE model tuned for fast inference, high-throughput reasoning, and coding tasks.
DeepInfra
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.1000
Output / 1M$0.2000
LanguageDeepSeek V4 Pro
deepseek-ai/DeepSeek-V4-Pro
DeepSeek V4 Pro is a 1M-token MoE model built for advanced reasoning, coding, and long-running agent tasks.
DeepInfra
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.3000
Output / 1M$2.6000
LanguageKimi K2.6
moonshotai/Kimi-K2.6
Moonshot AI's Kimi K2.6 model available through DeepInfra serverless inference.
DeepInfra
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.7500
Output / 1M$3.5000
LanguageMiMo-V2.5
XiaomiMiMo/MiMo-V2.5
MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding.
DeepInfra
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.4000
Output / 1M$2.0000
LanguageMiMo-V2.5-Pro
XiaomiMiMo/MiMo-V2.5-Pro
MiMo-V2.5-Pro is an open-source Mixture-of-Experts language model with 1.02T total parameters and 42B active parameters.
DeepInfra
1,048,576tokens
–Images✓JSON schema✓Function calling
Input / 1M$1.0000
Output / 1M$3.0000
LanguageQwen3.6-35B-A3B
Qwen/Qwen3.6-35B-A3B
Qwen3.6-35B-A3B is Alibaba's flagship Mixture-of-Experts model with 35B total parameters and only 3B activated per token.
DeepInfra
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.1500
Output / 1M$0.9500
LanguageMuse Spark 1.1
muse-spark-1.1
Meta's Muse Spark 1.1 model for agentic workflows, coding assistants, structured output, multimodal understanding, and long-context reasoning.
Meta
1,048,576tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.2500
Output / 1M$4.2500
LanguageKata 1.1 Medium
kata-1.1-medium
BotDojo-managed Kata 1.1 Medium model.
BotDojo
1,040,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.6100
Output / 1M$5.0600
LanguageKata 1.1 Low
kata-1.1-low
BotDojo-managed Kata 1.1 Low model.
BotDojo
512,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$0.3450
Output / 1M$1.3800
LanguageKata 1.0 High
kata-1.0-high
BotDojo-managed Kata 1.0 High model.
BotDojo
1,000,000tokens
✓Images✓JSON schema✓Function calling
Input / 1M$3.4500
Output / 1M$17.2500
Long context from 200,000 tokens
LanguageKata 1.0 Medium
kata-1.0-medium
BotDojo-managed Kata 1.0 Medium model.
BotDojo
262,144tokens
✓Images✓JSON schema✓Function calling
Input / 1M$1.0930
Output / 1M$4.6000
LanguageKata 1.0 Low
kata-1.0-low
BotDojo-managed Kata 1.0 Low model.
BotDojo
196,608tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3450
Output / 1M$1.3800
LanguageKata 1.0 Fast
kata-1.0-fast
BotDojo-managed Kata 1.0 Fast model.
BotDojo
196,608tokens
–Images✓JSON schema✓Function calling
Input / 1M$0.3450
Output / 1M$1.3800
EmbeddingText Embedding Ada 002
text-embedding-ada-002
Text Embedding Ada 002
OpenAI
8,191tokens1,536 dimensions
✓Embedding–Reduced dimensions
$0.1000per 1M tokens
EmbeddingText Embedding 3 Small
text-embedding-3-small
Increased performance over 2nd generation ada embedding model
OpenAI
8,191tokens1,536 dimensions
✓Embedding✓Reduced dimensions
$0.0200per 1M tokens
EmbeddingText Embedding 3 Large
text-embedding-3-large
Most capable embedding model for both english and non-english tasks
OpenAI
8,191tokens3,072 dimensions
✓Embedding✓Reduced dimensions
$0.1300per 1M tokens
EmbeddingEmbed English v3.0
embed-english-v3.0
A model that allows for text to be classified or turned into embeddings. English only.
Cohere
512tokens1,024 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingEmbed English Light v3.0
embed-english-light-v3.0
A smaller, faster version of embed-english-v3.0. Almost as capable, but a lot faster. English only.
Cohere
512tokens384 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingEmbed English v2.0
embed-english-v2.0
Our older embeddings model that allows for text to be classified or turned into embeddings. English only
Cohere
512tokens4,096 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingEmbed English Light v2.0
embed-english-light-v2.0
A smaller, faster version of embed-english-v2.0. Almost as capable, but a lot faster. English only.
Cohere
512tokens1,024 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingEmbed Multilingual v3.0
embed-multilingual-v3.0
Provides multilingual classification and embedding support. See supported languages here.
Cohere
512tokens1,024 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingEmbed Multilingual Light v3.0
embed-multilingual-light-v3.0
A smaller, faster version of embed-multilingual-v3.0. Almost as capable, but a lot faster. Supports multiple languages.
Cohere
512tokens384 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingEmbed Multilingual v2.0
embed-multilingual-v2.0
Provides multilingual classification and embedding support. See supported languages here.
Cohere
256tokens768 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingCohere Embed English
cohere.embed-english-v3
Cohere English Embedding Model hosted on AWS Bedrock
Amazon Bedrock
512tokens1,024 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingCohere Embed Multilingual
cohere.embed-multilingual-v3
Cohere Multilingual Embedding Model hosted on AWS Bedrock
Amazon Bedrock
512tokens1,024 dimensions
✓Embedding–Reduced dimensions
Not listed
EmbeddingAmazon Titan Embeddings G1 - Text
amazon.titan-embed-text-v1
Amazon's G1 Test Embedding Model hosted on AWS Bedrock
Amazon Bedrock
8,192tokens1,024 dimensions
✓Embedding–Reduced dimensions
$0.1000per 1M tokens
EmbeddingAmazon Titan Embeddings V2 - Text
amazon.titan-embed-text-v2:0
Amazon's G2 Text Embedding Model hosted on AWS Bedrock
Amazon Bedrock
8,192tokens1,024 dimensions
✓Embedding–Reduced dimensions
$0.1000per 1M tokens
EmbeddingOpenAI embedding Large
text-embedding-3-large
OpenAI's Large Text Embedding Model hosted on Microsoft Azure
Azure OpenAI
8,192tokens3,072 dimensions
✓Embedding✓Reduced dimensions
$0.1300per 1M tokens
EmbeddingOpenAI embedding Small
text-embedding-3-small
OpenAI's Small Text Embedding Model hosted on Microsoft Azure