Skip to main content

Model Pricing Directory

Live pricing for all available models on APIYI: input/output/cache prices, endpoints, and groups. Auto-updated.

Data updated at 2026-07-25 11:51 (UTC+8) from the live pricing API. All prices in USD. Top-ups settle at a fixed 1:7 exchange rate, stackable with recharge bonuses — see Pricing.

Models online

309

Vendors

16

Usage-based

273

Per-call

36

Discounts

Prices in the tables below are default list prices — your effective cost can be lower, since recharge promotions and group discounts stack:
  • Recharge promotions: with recharge bonuses your effective cost drops to 1/1.1 of list price (about 91%), up to 1/1.2 (about 83%) during campaigns — see Recharge promotions.
  • Group discounts: some billing groups carry extra discounts (e.g. 0.95×, i.e. 95% of list) — see Group overview.
The console reflects the actual charge in real time.

Other views

Besides this page, model pricing is also available at:

Column reference

  • Input / Output / Cache read: prices for usage-based models in USD per 1M tokens ($/1M). Cache read is the input price on prompt-cache hits; ”—” means no cache ratio is configured for the model.
  • Per-call: billed per request (image/video models), price in USD per call.
  • Endpoints: chat = /v1/chat/completions; responses = /v1/responses; messages = /v1/messages; gemini = native /v1beta API; images = /v1/images/generations; embeddings = /v1/embeddings.
  • Groups: billing groups where the model is available. Some groups carry discounts, stackable with recharge bonuses; the console is authoritative.
  • : tiered pricing — the rate depends on per-request token volume, see “Tiered pricing” at the end of the page.
  • This page is a complete single-page list — besides the search box above, desktop users can also use Ctrl+F (Cmd+F on macOS) to find a model name.

OpenAI

Anthropic

Google

xAI

DeepSeek

Alibaba

ByteDance

Moonshot

Zhipu

MiniMax

BAAI

Black Forest Labs

HappyHorse

Meituan

StepFun

Xiaomi

Tiered pricing

The following models are priced by per-request token volume (output price = tier input price × output multiplier):
  • gemini-2.5-pro: 0 – 200,000 tokens, input $1.25/1M; above 200,000 tokens, input $2.5/1M
  • gemini-3.1-pro-preview: 0 – 200,000 tokens, input $1.8/1M; above 200,000 tokens, input $3.6/1M (includes a standing discount — 90% of the nominal list tiers)
  • glm-5: 0 – 32,000 tokens, input $0.56/1M; above 32,000 tokens, input $0.86/1M
  • glm-5.1: 0 – 32,768 tokens, input $0.84/1M; above 32,768 tokens, input $1.14/1M
  • gpt-5.4: 0 – 278,528 tokens, input $2.5/1M; above 278,528 tokens, input $5/1M
  • gpt-5.4-pro: 0 – 278,528 tokens, input $30/1M; above 278,528 tokens, input $60/1M
  • gpt-5.5: 0 – 278,528 tokens, input $5/1M; above 278,528 tokens, input $10/1M
  • gpt-5.5-pro: 0 – 278,528 tokens, input $30/1M; above 278,528 tokens, input $60/1M
  • gpt-5.6-luna: 0 – 272,000 tokens, input $1/1M; above 272,000 tokens, input $2/1M
  • gpt-5.6-sol: 0 – 272,000 tokens, input $5/1M; above 272,000 tokens, input $10/1M
  • gpt-5.6-terra: 0 – 272,000 tokens, input $2.5/1M; above 272,000 tokens, input $5/1M
  • grok-4.3: 0 – 204,800 tokens, input $1.25/1M; above 204,800 tokens, input $2.5/1M
  • grok-4.5: 0 – 204,800 tokens, input $2/1M; above 204,800 tokens, input $4/1M
  • mimo-v2-pro: 0 – 262,144 tokens, input $1/1M; 262,145 – 1,024,000 tokens, input $2/1M
  • MiniMax-M3: 0 – 524,288 tokens, input $0.3/1M; above 524,288 tokens, input $0.6/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen-plus: 0 – 128,000 tokens, input $0.4/1M; 128,001 – 256,000 tokens, input $1.2/1M; 256,001 – 1,000,000 tokens, input $2.4/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen-plus-2025-09-11: 0 – 128,000 tokens, input $0.4/1M; 128,001 – 256,000 tokens, input $1.2/1M; 256,001 – 1,000,000 tokens, input $2.4/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen-plus-latest: 0 – 128,000 tokens, input $0.4/1M; 128,001 – 256,000 tokens, input $1.2/1M; 256,001 – 1,000,000 tokens, input $2.4/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen3-coder-480b-a35b-instruct: 0 – 32,768 tokens, input $3/1M; 32,769 – 131,072 tokens, input $5.4/1M; 131,073 – 204,800 tokens, input $9/1M
  • qwen3-coder-flash: 0 – 32,000 tokens, input $0.5/1M; 32,001 – 128,000 tokens, input $0.75/1M; 128,001 – 256,000 tokens, input $1.25/1M; 256,001 – 1,000,000 tokens, input $2.5/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen3-coder-plus: 0 – 32,000 tokens, input $5/1M; 32,001 – 128,000 tokens, input $7.5/1M; 128,001 – 256,000 tokens, input $12.5/1M; 256,001 – 1,000,000 tokens, input $25/1M
  • qwen3-coder-plus-2025-07-22: 0 – 32,000 tokens, input $2/1M; 32,001 – 128,000 tokens, input $3/1M; 128,001 – 256,000 tokens, input $5/1M; 256,001 – 1,000,000 tokens, input $10/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen3-coder-plus-2025-09-23: 0 – 32,000 tokens, input $2/1M; 32,001 – 128,000 tokens, input $3/1M; 128,001 – 256,000 tokens, input $5/1M; 256,001 – 1,000,000 tokens, input $10/1M (includes a standing discount — 50% of the nominal list tiers)
  • qwen3-max: 0 – 32,000 tokens, input $1.2/1M; 32,001 – 128,000 tokens, input $1.92/1M; 128,001 – 256,000 tokens, input $3.36/1M (includes a standing discount — 48% of the nominal list tiers)
  • qwen3-max-2025-09-23: 0 – 32,000 tokens, input $1.2/1M; 32,001 – 128,000 tokens, input $2/1M; 128,001 – 256,000 tokens, input $3/1M (includes a standing discount — 20% of the nominal list tiers)
  • qwen3-max-preview: 0 – 32,000 tokens, input $1.2/1M; 32,001 – 128,000 tokens, input $2/1M; 128,001 – 256,000 tokens, input $3/1M (includes a standing discount — 20% of the nominal list tiers)
  • qwen3-vl-plus: 0 – 32,000 tokens, input $0.3/1M; 32,001 – 128,000 tokens, input $0.45/1M; 128,001 – 256,000 tokens, input $0.9/1M (includes a standing discount — 30% of the nominal list tiers)
  • qwen3-vl-plus-2025-09-23: 0 – 32,000 tokens, input $0.3/1M; 32,001 – 128,000 tokens, input $0.45/1M; 128,001 – 256,000 tokens, input $0.9/1M (includes a standing discount — 30% of the nominal list tiers)
  • qwen3.5-flash: 0 – 131,072 tokens, input $0.03/1M; 131,073 – 262,144 tokens, input $0.122143/1M; above 262,144 tokens, input $0.182143/1M
  • qwen3.5-flash-2026-02-23: 0 – 131,072 tokens, input $0.03/1M; 131,073 – 262,144 tokens, input $0.122143/1M; above 262,144 tokens, input $0.182143/1M
  • qwen3.5-plus: 0 – 131,072 tokens, input $0.114/1M; 131,073 – 262,144 tokens, input $0.28/1M; above 262,144 tokens, input $0.56/1M
  • qwen3.5-plus-2026-02-15: 0 – 131,072 tokens, input $0.114/1M; 131,073 – 262,144 tokens, input $0.28/1M; above 262,144 tokens, input $0.56/1M
  • qwen3.6-flash: 0 – 262,144 tokens, input $0.17/1M; 262,145 – 1,024,000 tokens, input $0.68/1M
  • qwen3.6-max-preview: 0 – 131,072 tokens, input $1.28/1M; 131,073 – 262,144 tokens, input $2.12/1M
  • qwen3.6-plus: 0 – 262,144 tokens, input $0.3/1M; 262,145 – 1,024,000 tokens, input $1.2/1M
  • qwen3.7-plus: 0 – 256,000 tokens, input $0.6/1M; 256,001 – 1,000,000 tokens, input $1.8/1M (includes a standing discount — 30% of the nominal list tiers)
  • seed-2-0-lite-260228: 0 – 131,072 tokens, input $0.25/1M; 131,073 – 262,144 tokens, input $0.5/1M
  • seed-2-0-mini-260215: 0 – 131,072 tokens, input $0.1/1M; 131,073 – 262,144 tokens, input $0.2/1M
  • seed-2-0-pro-260328: 0 – 131,072 tokens, input $0.5/1M; 131,073 – 262,144 tokens, input $1/1M
Prices are subject to change; the console reflects live rates. See the API Capabilities section for per-model usage guides.