Models
Models available through the gateway, with current pricing.
Anthropic’s newest flagship generation with a 1M-token context — deep multi-step reasoning, agentic autonomy and long-session coherence beyond Opus 4.6.
Fast, inexpensive Claude model for high-volume tasks: classification, extraction, summarisation and latency-sensitive chat.
Opus-tier Claude model for the hardest reasoning, agentic and coding workloads.
Anthropic’s most capable 4.x model for deep reasoning, agentic workflows and long-horizon coding tasks, with 1M context.
Opus-tier Claude model for the hardest reasoning, agentic and coding workloads.
Opus-tier Claude model for the hardest reasoning, agentic and coding workloads.
Next-generation Opus with a major jump in planning, tool orchestration and hard-problem reasoning. The premium tier of the Claude family.
Sonnet-tier Claude model balancing capability and cost for production workloads.
Balanced Claude model for production workloads — near-Opus quality on everyday tasks at a fraction of the cost, with strong tool use.
Fifth-generation Sonnet: near-Opus capability at $2/$10 — the new price-performance sweet spot of the Claude family.
DeepSeek general chat model — remarkable quality-per-dollar for chat, RAG and structured extraction.
DeepSeek coding model for completion, generation and repair.
DeepSeek reasoning model exposing chain-of-thought at very low cost.
DeepSeek reasoning model exposing chain-of-thought for math, code and analysis at very low cost.
DeepSeek model with remarkable quality-per-dollar for chat and RAG.
DeepSeek model with remarkable quality-per-dollar for chat and RAG.
DeepSeek model with remarkable quality-per-dollar for chat and RAG.
xAI Grok reasoning model for coding, analysis and tool-driven agent workflows.
xAI Grok reasoning model for coding, analysis and tool-driven agent workflows.
Moonshot Kimi open-weights MoE model with strong agentic and multilingual performance.
Moonshot Kimi reasoning model interleaving step-by-step thinking with tool calls.
Moonshot Kimi reasoning model interleaving step-by-step thinking with tool calls.
High-throughput turbo tier of Moonshot Kimi for latency-sensitive workloads.
Moonshot’s open-weights MoE flagship with strong agentic and multilingual performance.
Moonshot Kimi open-weights MoE model with strong agentic and multilingual performance.
Moonshot Kimi coding model tuned for agentic software engineering.
Moonshot’s third-generation flagship MoE — frontier-level agentic and reasoning performance at aggressive open pricing.
Moonshot Kimi open-weights MoE model with strong agentic and multilingual performance.
General-purpose GPT model with strong reasoning and structured output.
Codex variant tuned for agentic software engineering and long coding sessions.
Cost-efficient GPT tier for everyday chat, RAG and high-throughput workloads.
Cost-efficient GPT tier for everyday chat, RAG and high-throughput workloads.
Pro tier of the GPT line for maximum quality on the hardest problems.
General-purpose GPT model with strong reasoning and structured output.
Codex variant tuned for agentic software engineering and long coding sessions.
Codex variant tuned for agentic software engineering and long coding sessions.
OpenAI’s flagship general-purpose model with strong reasoning, multimodal understanding and reliable structured output.
Agentic coding variant of GPT-5.2 tuned for long-running software engineering sessions, refactors and code review.
Cost-efficient GPT tier for everyday chat, RAG and high-throughput workloads.
Pro tier of the GPT line for maximum quality on the hardest problems.
Codex variant tuned for agentic software engineering and long coding sessions.
General-purpose GPT model with strong reasoning and structured output.
Cost-efficient GPT tier for everyday chat, RAG and high-throughput workloads.
Cost-efficient GPT tier for everyday chat, RAG and high-throughput workloads.
Pro tier of the GPT line for maximum quality on the hardest problems.
General-purpose GPT model with strong reasoning and structured output.
Pro tier of the GPT line for maximum quality on the hardest problems.
Latest GPT generation with a 1M-token window; the strongest general model in the GPT line-up.
Cost-efficient GPT tier for everyday chat, RAG and high-throughput workloads.
General-purpose GPT model with strong reasoning and structured output.
General-purpose GPT model with strong reasoning and structured output.
OpenAI image-generation model for high-quality text-to-image and image editing, priced per generated image by output resolution.
Zhipu GLM model with strong coding and agent capabilities.
Lightweight GLM tier for fast, cheap, high-volume completions.
Zhipu’s workhorse GLM model — a popular Claude-compatible daily driver for coding and agents.
Zhipu GLM model with strong coding and agent capabilities.
Lightweight GLM tier for fast, cheap, high-volume completions.
Zhipu’s fifth-generation flagship with strong coding and agent capabilities.
Zhipu GLM coding model tuned for software engineering and agent use.
Zhipu GLM model with strong coding and agent capabilities.
Zhipu GLM model with strong coding and agent capabilities.
Zhipu GLM model with strong coding and agent capabilities.