The fastest Claude, near-frontier intelligence — 200k context, extended thinking.
Model catalog
26 frontier models. All live today.
Grouped by provider: Anthropic, OpenAI (+ GPT Image 2), Google (Gemini + Nano Banana), DeepSeek, Doubao, GLM, Kimi, MiniMax. One key, one base URL — official list price next to our price on each, pulled live from the gateway catalog.
Anthropic
7Previous-gen flagship — 1M context, extended thinking. Consider migrating to Opus 4.7.
Anthropic's most capable general-purpose model — excels at complex reasoning and agentic coding, native 1M context.
Anthropic's newest flagship — further gains in complex reasoning and agentic coding, native 1M context.
Anthropic's newest flagship Claude Opus 5 — leading complex reasoning and agentic coding, native 1M context.
The best balance of speed and intelligence — 1M context, extended/adaptive thinking.
The best balance of speed and intelligence — adaptive thinking on by default, native 1M context.
OpenAI
8OpenAI's more economical coding and professional-work model — 1M context.
OpenAI's strongest mini model — good for coding, computer use, and subagents; 400k context.
OpenAI's flagship coding and professional-work model — adjustable reasoning effort, 1M context.
Official GPT-5.6 alias — always routes to the Sol flagship tier.
OpenAI's cost-sensitive high-throughput tier (Luna) — extraction, classification and summarization workloads, 1.05M context.
OpenAI's frontier flagship (Sol tier) — complex coding and long-horizon agentic work, reasoning effort up to max, 1.05M context.
OpenAI's balanced intelligence/cost tier (Terra) — GPT-5.5-class performance at a lower price, 1.05M context.
OpenAI's native image generation and editing (/v1/images generations + edits) — supports arbitrary resolutions.
Low-latency, cost-efficient multimodal model — good for translation, transcription, document processing, and other high-frequency lightweight tasks.
Google's most intelligent model — sustained frontier performance, excels at agentic and coding tasks.
Always points to the newest Gemini Flash generation; currently equivalent to gemini-3.5-flash.
Always points to the newest Gemini Flash-Lite generation; currently equivalent to gemini-3.1-flash-lite.
Google's flagship thinking model — excels at complex code, math, and STEM reasoning; 1M context.
DeepSeek
2DeepSeek's faster, more economical API. Official DeepSeek traffic is served by the vision SKU (image understanding). Deep thinking on by default, can be disabled.
DeepSeek V4 Pro — significantly stronger agent capability and rich world knowledge; deep thinking on by default, can be disabled manually.
GLM
1Zhipu's first native multimodal model in the GLM-5 series — 320B total / 18B active parameters, with image/video/file input and 1M-token context.
Kimi
1Kimi's flagship multimodal model (text/image/video input) with a 1M-token context window and always-on reasoning; Kimi Code defaults reasoning_effort to high.
More
2xAI's intelligent coding model for agentic software and engineering workflows — 500k context, tunable reasoning_effort (default high).
xAI's flagship model for coding, long-horizon agents, and knowledge work — 500k context, tunable reasoning_effort (default high, including xhigh).
Need a model that isn't here? Email sales@bytespike.ai — most new model releases land on the gateway within the same week.