$5
Pricing
Official price vs. our price, on every model.
26 frontier models from 8 providers behind one key — Anthropic, OpenAI, Google, DeepSeek, Doubao, GLM, Kimi, MiniMax. Claude as low as 82% off; DeepSeek · Doubao · GLM · Kimi at half price; OpenAI text & Gemini at list price. Priced in plain USD. Top up online or activate with a redeem code. Failures don't bill.
Discounts vs. official
Frontier models, well under list price.
We publish each provider's official list price next to ours — the tiers below are read live from the gateway rate card, so they always match what you're billed.
How to activate
Ways to get started today.
Top up online in the Console in seconds. Prefer another route? You can also:
Top up
Priced in plain US dollars.
Your balance never expires. Bonus tiers kick in automatically — the dollar amount you top up determines the bonus.
$50
≈ $50.75 effective
Top up$500
≈ $525 effective
Top up$1,250
≈ $1,390 effective
Top upYour balance never expires. Bonuses are applied at top-up time, not on spend. Failed requests don't bill — preview the cost via the x-bytespike-credits-estimated response header before the request runs.
Per-endpoint pricing
Official price vs. our price, per model.
Every provider's official list price next to ours, grouped by brand — discounts pulled live from the gateway's published rate card. Failures don't bill.
Anthropic
7| Model | Official | Our price | Discount |
|---|---|---|---|
| Claude Haiku 4.5 The fastest Claude, near-frontier intelligence — 200k context, extended thinking. | $1 / $5 | $0.18 / $0.9per 1M tokens | 82% off |
| Claude Opus 4.6 Previous-gen flagship — 1M context, extended thinking. Consider migrating to Opus 4.7. | $5 / $25 | $0.9 / $4.5per 1M tokens | 82% off |
| Claude Opus 4.7 Anthropic's most capable general-purpose model — excels at complex reasoning and agentic coding, native 1M context. | $5 / $25 | $0.9 / $4.5per 1M tokens | 82% off |
| Claude Opus 4.8 Anthropic's newest flagship — further gains in complex reasoning and agentic coding, native 1M context. | $5 / $25 | $0.9 / $4.5per 1M tokens | 82% off |
| Claude Opus 5 Anthropic's newest flagship Claude Opus 5 — leading complex reasoning and agentic coding, native 1M context. | $5 / $25 | $0.9 / $4.5per 1M tokens | 82% off |
| Claude Sonnet 4.6 The best balance of speed and intelligence — 1M context, extended/adaptive thinking. | $3 / $15 | $0.54 / $2.7per 1M tokens | 82% off |
| Claude Sonnet 5 The best balance of speed and intelligence — adaptive thinking on by default, native 1M context. | $2 / $10 | $0.36 / $1.8per 1M tokens | 82% off |
OpenAI
8| Model | Official | Our price | Discount |
|---|---|---|---|
| GPT-5.4 OpenAI's more economical coding and professional-work model — 1M context. | $2.5 / $15 | $2 / $12per 1M tokens | 20% off |
| GPT-5.4 mini OpenAI's strongest mini model — good for coding, computer use, and subagents; 400k context. | $0.75 / $4.5 | $0.6 / $3.6per 1M tokens | 20% off |
| GPT-5.5 OpenAI's flagship coding and professional-work model — adjustable reasoning effort, 1M context. | $5 / $30 | $4 / $24per 1M tokens | 20% off |
| GPT-5.6 (Sol) Official GPT-5.6 alias — always routes to the Sol flagship tier. | $4 / $20 | $3.2 / $16per 1M tokens | 20% off |
| GPT-5.6 Luna OpenAI's cost-sensitive high-throughput tier (Luna) — extraction, classification and summarization workloads, 1.05M context. | $0.2 / $1.2 | $0.16 / $0.96per 1M tokens | 20% off |
| GPT-5.6 Sol OpenAI's frontier flagship (Sol tier) — complex coding and long-horizon agentic work, reasoning effort up to max, 1.05M context. | $5 / $30 | $4 / $24per 1M tokens | 20% off |
| GPT-5.6 Terra OpenAI's balanced intelligence/cost tier (Terra) — GPT-5.5-class performance at a lower price, 1.05M context. | $2 / $12 | $1.6 / $9.6per 1M tokens | 20% off |
| GPT Image 2 OpenAI's native image generation and editing (/v1/images generations + edits) — supports arbitrary resolutions. | $5 / $10 | $4 / $8per 1M tokens | 20% off |
| Model | Official | Our price | Discount |
|---|---|---|---|
| Gemini 3.1 Flash-Lite Low-latency, cost-efficient multimodal model — good for translation, transcription, document processing, and other high-frequency lightweight tasks. | $0.25 / $1.5 | $0.25 / $1.5per 1M tokens | List price |
| Gemini 3.5 Flash Google's most intelligent model — sustained frontier performance, excels at agentic and coding tasks. | $1.5 / $9 | $1.5 / $9per 1M tokens | List price |
| Gemini Flash (latest) Always points to the newest Gemini Flash generation; currently equivalent to gemini-3.5-flash. | $0.3 / $2.5 | $0.3 / $2.5per 1M tokens | List price |
| Gemini Flash-Lite (latest) Always points to the newest Gemini Flash-Lite generation; currently equivalent to gemini-3.1-flash-lite. | $0.25 / $1.5 | $0.25 / $1.5per 1M tokens | List price |
| Gemini Pro (latest) Google's flagship thinking model — excels at complex code, math, and STEM reasoning; 1M context. | $2 / $12 | $2 / $12per 1M tokens | List price |
DeepSeek
2| Model | Official | Our price | Discount |
|---|---|---|---|
| DeepSeek V4 Flash DeepSeek's faster, more economical API. Official DeepSeek traffic is served by the vision SKU (image understanding). Deep thinking on by default, can be disabled. | $0.44 / $1.32 | $0.44 / $1.32per 1M tokens | List price |
| DeepSeek V4 Pro DeepSeek V4 Pro — significantly stronger agent capability and rich world knowledge; deep thinking on by default, can be disabled manually. | $1.32 / $3.96 | $1.32 / $3.96per 1M tokens | List price |
GLM
1| Model | Official | Our price | Discount |
|---|---|---|---|
| GLM 5.3 Flash Zhipu's first native multimodal model in the GLM-5 series — 320B total / 18B active parameters, with image/video/file input and 1M-token context. | $0.15 / $0.5 | $0.15 / $0.5per 1M tokens | List price |
Kimi
1| Model | Official | Our price | Discount |
|---|---|---|---|
| Kimi K3 Kimi's flagship multimodal model (text/image/video input) with a 1M-token context window and always-on reasoning; Kimi Code defaults reasoning_effort to high. | $3 / $15 | $3 / $15per 1M tokens | List price |
More
2| Model | Official | Our price | Discount |
|---|---|---|---|
| Grok 4.5 xAI's intelligent coding model for agentic software and engineering workflows — 500k context, tunable reasoning_effort (default high). | $2 / $6 | $2 / $6per 1M tokens | List price |
| Grok 4.6 xAI's flagship model for coding, long-horizon agents, and knowledge work — 500k context, tunable reasoning_effort (default high, including xhigh). | $2 / $6 | $2 / $6per 1M tokens | List price |
Token prices are per 1M (input / output). Official is each provider's public list price; our price applies the family discount. Image models are billed per token; video is billed per clip. Failures don't bill — the final charge matches the estimated_credits returned at submit time.
rate card refreshed · 2026-08-29
Per-model rate card
26 models, transparent prices, refreshed live.
We publish official list price next to our price per million tokens for every chat model, and per-call rates for image and video. Sourced from the production gateway, refreshed once a day. No hidden tiers, no markup paragraphs.
Chat models
25
Image models
1
One key, one URL
llm.bytespike.ai/v1
All prices are quoted in U.S. dollars (USD). ByteSpike is pay-as-you-go: top up your balance, spend per the rate card — no plans, no minimums, no auto-renewal. Top up online at console.bytespike.ai/billing, or use a redeem code (console.bytespike.ai/billing#redeem). Your top-up balance never expires and failed requests don't bill. Stated prices exclude applicable taxes and VAT, calculated at activation based on your billing location.