// all accessible via one api key

Model Catalog

Browse every model available through LLMAI. Filter by provider, compare context windows and capabilities, and copy the slug directly into your code.

GOpenAI
G
GPT-4.1NEW
gpt-4.1
1M ctx

Reliable multimodal model for coding, instruction following, tool use, and long-context applications.

ChatCodeVision
Input / 1M tokens
$1.000
Output / 1M tokens
$4.000
G
GPT-4oNEW
gpt-4o
128K ctx

Fast multimodal model for interactive chat, vision, extraction, and general application workloads.

ChatVisionFast
Input / 1M tokens
$1.250
Output / 1M tokens
$5.000
G
GPT-5NEW
gpt-5
400K ctx

General-purpose GPT-5 reasoning model for complex analysis, coding, and agentic workflows.

ChatReasoningCodeVision
Input / 1M tokens
$0.625
Output / 1M tokens
$5.000
G
GPT-5 MiniNEW
gpt-5-mini
400K ctx

Efficient GPT-5 model for high-volume reasoning, coding assistance, and structured tasks.

ChatReasoningFastVision
Input / 1M tokens
$0.125
Output / 1M tokens
$1.000
G
GPT-5.4
gpt-5.4
1M ctx

Mid-tier GPT-5 with strong general-purpose reasoning, multimodal input, and tool-calling support — the entry GPT-5 tier.

ChatReasoningVision
Input / 1M tokens
$0.787
Output / 1M tokens
$5.250
G
GPT-5.5
gpt-5.5
1M ctx

Flagship GPT-5.5. The most capable OpenAI model on the platform — deep reasoning, long-context comprehension, and frontier multimodal performance.

ChatReasoningVisionLong Context
Input / 1M tokens
$1.575
Output / 1M tokens
$10.500
G
GPT-5.6 Sol
gpt-5.6-sol
1M ctx

Complex professional work, reasoning and coding. The flagship GPT-5.6 model — strongest reasoning and code performance in the family.

ChatReasoningCodingVision
Input / 1M tokens
$2.500
Output / 1M tokens
$15.000
G
GPT-5.6 Terra
gpt-5.6-terra
1M ctx

A balanced option for intelligence and cost. Solid general-purpose reasoning at a lower price than Sol.

ChatReasoningVision
Input / 1M tokens
$1.250
Output / 1M tokens
$7.500
G
GPT-5.6 Luna
gpt-5.6-luna
1M ctx

Optimized for cost-sensitive, high-volume workloads. The cheapest GPT-5.6 model for fast everyday tasks.

ChatVision
Input / 1M tokens
$0.500
Output / 1M tokens
$3.000
AAnthropic
A
Claude Opus 5NEW
claude-opus-5
1M ctx

Anthropic's latest Opus model. Built for frontier coding, agentic reasoning, and complex long-horizon work.

ChatReasoningCodeVision
Input / 1M tokens
$2.625
Output / 1M tokens
$10.500
A
Claude Opus 4.8
claude-opus-4.8
1M ctx

Proven frontier Opus model. Excellent for complex multi-step engineering, research synthesis, and high-stakes reasoning.

ChatReasoningCodeVision
Input / 1M tokens
$2.625
Output / 1M tokens
$10.500
A
Claude Opus 4.7
claude-opus-4.7
1M ctx

Previous-generation flagship Opus. Excellent for complex multi-step engineering, research synthesis, and high-stakes reasoning.

ChatReasoningCodeVision
Input / 1M tokens
$2.625
Output / 1M tokens
$10.500
A
Claude Opus 4.6
claude-opus-4.6
1M ctx

The cost-efficient Opus tier. Strong reasoning at a meaningfully lower per-token price than 4.7/4.8.

ChatReasoningCodeVision
Input / 1M tokens
$1.575
Output / 1M tokens
$7.875
A
Claude Sonnet 4.6
claude-sonnet-4.6
1M ctx

Balanced Sonnet — high-volume daily-driver for chat, code, and structured output. The recommended starting point in the Claude family.

ChatCodeVision
Input / 1M tokens
$1.050
Output / 1M tokens
$5.250
A
Claude Sonnet 5NEW
claude-sonnet-5
1M ctx

Anthropic's newest Sonnet — upgraded reasoning, code, and agentic performance over Sonnet 4.6 at the same price. The new default Claude for daily work.

ChatReasoningCodeVision
Input / 1M tokens
$1.050
Output / 1M tokens
$5.250
A
Claude Haiku 4.5
claude-haiku-4.5
1M ctx

Anthropic's fastest and most cost-efficient Claude model. Ideal for high-volume chat, classification, summarisation, and lightweight agent loops where latency and cost matter most.

ChatFastVision
Input / 1M tokens
$0.525
Output / 1M tokens
$2.625
Google
Gemini 3.5 FlashNEW
gemini-3.5-flash
1M ctx

Google's latest production Flash model, with frontier intelligence, native multimodal input, search grounding, and a 1M-token context window.

ChatFastReasoningVisionLong Context
Input / 1M tokens
$0.750
Output / 1M tokens
$4.500
Gemini 3.1 Flash LiteNEW
gemini-3.1-flash-lite
1M ctx

Google's cost-efficient production model for high-volume agentic tasks, translation, classification, and data processing.

ChatFastReasoningVisionLong Context
Input / 1M tokens
$0.125
Output / 1M tokens
$0.750
Gemini 3 Flash (Preview)
gemini-3-flash-preview
1M ctx

A fast multimodal preview model with strong reasoning, search grounding, and a 1M-token context window.

ChatFastReasoningVisionLong Context
Input / 1M tokens
$0.250
Output / 1M tokens
$1.500
Gemini 2.5 ProNEW
gemini-2.5-pro
1M ctx

Google's production model for advanced coding, complex reasoning, and long-context multimodal work.

ChatReasoningCodeVisionLong Context
Input / 1M tokens
$0.625
Output / 1M tokens
$5.000
Gemini 2.5 Flash
gemini-2.5-flash
1M ctx

Google's production multimodal Flash model with reasoning support and a 1,048,576-token context window.

ChatReasoningVisionLong Context
Input / 1M tokens
$0.150
Output / 1M tokens
$1.250
Gemini 2.5 Flash Lite
gemini-2.5-flash-lite
1M ctx

The lowest-cost Gemini 2.5 tier for high-volume multimodal classification, extraction, and chat.

ChatFastReasoningVisionLong Context
Input / 1M tokens
$0.050
Output / 1M tokens
$0.200
Gemma 4
gemma-4
128K ctx

Google's open-weight Gemma 4 (31B) served via Ollama Cloud. Excellent value for everyday chat and code-completion tasks.

ChatFastCode
Input / 1M tokens
$0.046
Output / 1M tokens
$0.130
Veo 3.1
veo-3.1
ctx

Google's text-to-video generation model. Produces short, high-quality video clips from natural-language prompts.

Video
Per generation
$0.21
Flat fee per successful video — no token billing.
XxAI
X
Grok 4.5
grok-4.5
500K ctx

xAI's frontier reasoning model for coding, knowledge work, and STEM tasks with a 500K-token context window.

ChatReasoningCodeLong Context
Input / 1M tokens
$1.000
Output / 1M tokens
$3.000
KKimi
K
Kimi K3NEW
kimi-k3
1M ctx

Moonshot AI's flagship 2.8T-parameter model for long-horizon software engineering, end-to-end knowledge work, native visual understanding, and deep reasoning. Supports a 1M-token context window and configurable reasoning effort.

ChatReasoningCodeVisionLong Context
Input / 1M tokens
$1.500
Output / 1M tokens
$7.500
K
Kimi K2.7 Code
kimi-k2.7-code
256K ctx

Moonshot AI's code-specialized K2.7 variant. Tuned for code generation, refactoring, and agentic coding workflows — with native long thinking and deep reasoning.

ChatReasoningCodeLong Context
Input / 1M tokens
$0.665
Output / 1M tokens
$2.800
K
Kimi K2.6
kimi-k2.6
256K ctx

Moonshot AI's multimodal K2.6 model for dialogue and agent tasks, with visual input plus optional thinking and non-thinking modes.

ChatReasoningLong Context
Input / 1M tokens
$0.620
Output / 1M tokens
$2.600
K
Kimi K2.5
kimi-k2.5
256K ctx

The cost-efficient predecessor to K2.6, retaining strong long-context handling at a fraction of the price.

ChatReasoningLong Context
Input / 1M tokens
$0.210
Output / 1M tokens
$1.050
DeepSeek
DeepSeek R1NEW
deepseek-r1
128K ctx

Reasoning-focused model for mathematics, analysis, coding, and multi-step problem solving.

ChatReasoningCode
Input / 1M tokens
$0.275
Output / 1M tokens
$1.095
DeepSeek V4 Pro
deepseek-v4-pro
128K ctx

DeepSeek's strongest model. Tuned for advanced coding tasks, technical reasoning, and structured output generation.

ChatCodeReasoning
Input / 1M tokens
$0.217
Output / 1M tokens
$0.435
MMiniMax
M
MiniMax M3
minimax-m3
1M ctx

MiniMax's flagship long-context model with a 1M-token window. Ideal for whole-codebase analysis and large-document tasks.

ChatReasoningLong Context
Input / 1M tokens
$0.420
Output / 1M tokens
$1.680
M
MiniMax M2.7
minimax-m2.7
1M ctx

The efficient long-context model. 1M-token window at a budget price — great for retrieval-heavy and document workflows.

ChatLong ContextFast
Input / 1M tokens
$0.200
Output / 1M tokens
$0.850
M
MiniMax M2.5
minimax-m2.5
200K ctx

A cost-efficient MiniMax reasoning model for coding and agent workflows with a 200K-token context window.

ChatReasoningCodeLong Context
Input / 1M tokens
$0.150
Output / 1M tokens
$0.600
AAlibaba
A
Qwen 3.7 Max
qwen3.7-max
1M ctx

Alibaba's flagship Qwen 3.7 model for complex reasoning, coding, and long-context agent workflows.

ChatReasoningCodeLong Context
Input / 1M tokens
$1.250
Output / 1M tokens
$3.750
A
Qwen 3.7 Plus
qwen3.7-plus
1M ctx

The cost-efficient Qwen 3.7 tier with a 1M-token context window for general-purpose agents and coding.

ChatReasoningCodeLong Context
Input / 1M tokens
$0.200
Output / 1M tokens
$0.800
A
Qwen 3.5 Plus
qwen3.5-plus
128K ctx

The previous-generation Qwen-Plus model. Solid baseline for general chat and structured tasks.

ChatReasoningCode
Input / 1M tokens
$0.140
Output / 1M tokens
$0.810
XXiaomi
X
MiMo V2.5 Pro
mimo-v2.5-pro
128K ctx

Xiaomi's flagship MiMo model. Tuned for complex agent workflows, code-focused tasks, and structured reasoning.

ChatReasoningCode
Input / 1M tokens
$0.304
Output / 1M tokens
$0.609
X
MiMo V2.5
mimo-v2.5
128K ctx

The cost-efficient MiMo tier. Quick, capable, and great for high-volume general-purpose calls.

ChatFastReasoning
Input / 1M tokens
$0.098
Output / 1M tokens
$0.196
ZZ.AI
Z
GLM 5.3 Flash
glm-5.3-flash
1M ctx

Z.AI's efficient multimodal GLM model for coding, agentic work, and visual understanding with a 1M-token context window. Thinking is always enabled.

ChatFastReasoningCodeVisionLong Context
Input / 1M tokens
$0.045
Cached / 1M
$0.009
Output / 1M tokens
$0.150
Z
GLM 5
glm-5
200K ctx

Z.AI's GLM 5 reasoning model for coding and agentic tasks with a 200K-token context window.

ChatReasoningCodeLong Context
Input / 1M tokens
$0.500
Output / 1M tokens
$1.600
Z
GLM 5.2
glm-5.2
128K ctx

Z.AI's flagship GLM model for long-horizon agentic tasks. Served via Ollama Cloud with native prefix caching.

ChatReasoningCode
Input / 1M tokens
$0.700
Output / 1M tokens
$2.200
Z
GLM 5.1
glm-5.1
128K ctx

The latest iteration of Z.AI's GLM family, served via Ollama Cloud. Strong instruction-following and bilingual reasoning.

ChatReasoningCode
Input / 1M tokens
$0.550
Output / 1M tokens
$1.300