Model Catalog
Browse every model available through LLMAI. Filter by provider, compare context windows and capabilities, and copy the slug directly into your code.
Reliable multimodal model for coding, instruction following, tool use, and long-context applications.
Fast multimodal model for interactive chat, vision, extraction, and general application workloads.
General-purpose GPT-5 reasoning model for complex analysis, coding, and agentic workflows.
Efficient GPT-5 model for high-volume reasoning, coding assistance, and structured tasks.
Mid-tier GPT-5 with strong general-purpose reasoning, multimodal input, and tool-calling support — the entry GPT-5 tier.
Flagship GPT-5.5. The most capable OpenAI model on the platform — deep reasoning, long-context comprehension, and frontier multimodal performance.
Complex professional work, reasoning and coding. The flagship GPT-5.6 model — strongest reasoning and code performance in the family.
A balanced option for intelligence and cost. Solid general-purpose reasoning at a lower price than Sol.
Optimized for cost-sensitive, high-volume workloads. The cheapest GPT-5.6 model for fast everyday tasks.
Anthropic's latest Opus model. Built for frontier coding, agentic reasoning, and complex long-horizon work.
Proven frontier Opus model. Excellent for complex multi-step engineering, research synthesis, and high-stakes reasoning.
Previous-generation flagship Opus. Excellent for complex multi-step engineering, research synthesis, and high-stakes reasoning.
The cost-efficient Opus tier. Strong reasoning at a meaningfully lower per-token price than 4.7/4.8.
Balanced Sonnet — high-volume daily-driver for chat, code, and structured output. The recommended starting point in the Claude family.
Anthropic's newest Sonnet — upgraded reasoning, code, and agentic performance over Sonnet 4.6 at the same price. The new default Claude for daily work.
Anthropic's fastest and most cost-efficient Claude model. Ideal for high-volume chat, classification, summarisation, and lightweight agent loops where latency and cost matter most.
Google's latest production Flash model, with frontier intelligence, native multimodal input, search grounding, and a 1M-token context window.
Google's cost-efficient production model for high-volume agentic tasks, translation, classification, and data processing.
A fast multimodal preview model with strong reasoning, search grounding, and a 1M-token context window.
Google's production model for advanced coding, complex reasoning, and long-context multimodal work.
Google's production multimodal Flash model with reasoning support and a 1,048,576-token context window.
The lowest-cost Gemini 2.5 tier for high-volume multimodal classification, extraction, and chat.
Google's open-weight Gemma 4 (31B) served via Ollama Cloud. Excellent value for everyday chat and code-completion tasks.
Google's text-to-video generation model. Produces short, high-quality video clips from natural-language prompts.
xAI's frontier reasoning model for coding, knowledge work, and STEM tasks with a 500K-token context window.
Moonshot AI's flagship 2.8T-parameter model for long-horizon software engineering, end-to-end knowledge work, native visual understanding, and deep reasoning. Supports a 1M-token context window and configurable reasoning effort.
Moonshot AI's code-specialized K2.7 variant. Tuned for code generation, refactoring, and agentic coding workflows — with native long thinking and deep reasoning.
Moonshot AI's multimodal K2.6 model for dialogue and agent tasks, with visual input plus optional thinking and non-thinking modes.
The cost-efficient predecessor to K2.6, retaining strong long-context handling at a fraction of the price.
Reasoning-focused model for mathematics, analysis, coding, and multi-step problem solving.
DeepSeek's strongest model. Tuned for advanced coding tasks, technical reasoning, and structured output generation.
MiniMax's flagship long-context model with a 1M-token window. Ideal for whole-codebase analysis and large-document tasks.
The efficient long-context model. 1M-token window at a budget price — great for retrieval-heavy and document workflows.
A cost-efficient MiniMax reasoning model for coding and agent workflows with a 200K-token context window.
Alibaba's flagship Qwen 3.7 model for complex reasoning, coding, and long-context agent workflows.
The cost-efficient Qwen 3.7 tier with a 1M-token context window for general-purpose agents and coding.
The previous-generation Qwen-Plus model. Solid baseline for general chat and structured tasks.
Xiaomi's flagship MiMo model. Tuned for complex agent workflows, code-focused tasks, and structured reasoning.
The cost-efficient MiMo tier. Quick, capable, and great for high-volume general-purpose calls.
Z.AI's efficient multimodal GLM model for coding, agentic work, and visual understanding with a 1M-token context window. Thinking is always enabled.
Z.AI's GLM 5 reasoning model for coding and agentic tasks with a 200K-token context window.
Z.AI's flagship GLM model for long-horizon agentic tasks. Served via Ollama Cloud with native prefix caching.
The latest iteration of Z.AI's GLM family, served via Ollama Cloud. Strong instruction-following and bilingual reasoning.