Skip to main content

Overview

doubao-seedance-2-0-260128 (standard), doubao-seedance-2-0-fast-260128 (fast), and doubao-seedance-2-0-mini-260615 (mini/lite) are ByteDance’s latest video generation model family — three models running in parallel, served through APIYI on official Volcengine Mainland China resources (not the BytePlus international edition) with upstream content-safety built in. They support text-to-video, first+last/first frame image-to-video, and multi-modal inputs (0-9 reference images + 0-3 reference videos / 0-3 reference audios) — and can generate voice, sound effects, and background music synchronized with the visuals. Mini, added in June 2026, is the cost-efficiency pick: about half the standard model’s unit price and faster generation, capped at 720p.
🎬 Highlights: 4-15 s controllable duration (or -1 for model-chosen length), three resolution tiers (480p/720p/1080p; 1080p standard model only), 6 aspect ratios plus adaptive, synchronized audio on by default, and multilingual prompts (Chinese, English, Japanese, Spanish, Portuguese, Indonesian). Built for short-video production, e-commerce assets, motion design, and virtual-human content at scale.

Video Generation API Reference

POST /seedance/api/v3/contents/generations/tasks — async task endpoint with an interactive Playground and full polling/download code.

API Manual

Token creation, base URL, billing models, and general calling conventions.

Visual API Testing

Debug this endpoint directly in the iCover visual testing tool — no code required.

Async Task Lookup / Download

View submitted video tasks and download video links in the APIYI console — a lookup entry outside the API.

Why APIYI’s Seedance 2.0?

A note on positioning first: this model carries no official discount, and APIYI doesn’t price it for profit — it is offered to secure supply and serve customers. The real value of going through APIYI is not “cheaper”, but access and experience:

Official Resource · Mainland Edition

Official Volcengine Mainland China resources (not BytePlus international), with upstream content safety built in. Parameters, responses, and billing match the official API exactly.

Unlimited Concurrency · No Queuing

In our tests, 15 simultaneous tasks all entered running immediately with zero queuing (measured 2026-06-06 (UTC+8)) — ready for batch production at scale.

Supply-first Pricing · On Par with Official

No official discount exists, and APIYI doesn’t profit from this model: unit prices align with Volcengine’s official list (on-platform billing runs roughly 10% higher); combined with top-up bonuses the effective cost is about on par with the official channel, and high-tier recharge customers can land below it on some tiers.

Zero-friction Access · No ID Verification

No Volcengine account, no real-name/identity verification, no spending threshold (skip the CNY 200 activation deposit and enterprise verification). Mainland China data centers, residential networks, and overseas nodes can all reach api.apiyi.com directly with a single Token.

Virtual-face Whitelist Access

The channel comes with upstream virtual-face whitelist access: AI-generated faces and virtual avatars can be used directly for image-to-video, no separate whitelist application to the official channel needed (real human faces remain restricted by upstream content safety).

Full Video Model Lineup

VEO 3.1, Sora 2, and Wan2.7 are available on the same platform — mix and match per use case.

Professional Support

A team experienced in video-generation workloads, providing model selection, tuning, and integration support from PoC to production.

Key Features

Three Tiers · Same Price per Tier

480p / 720p / 1080p (1080p standard model only). Within a tier, 16:9, 9:16, 1:1, and every other ratio share the same pixel area and the same price — switch between landscape and portrait at zero cost.

Synchronized Audio by Default

generate_audio defaults to true: voice, sound effects, and background music are generated to match the visuals. Put spoken lines in double quotes to improve voice-over quality.

4-15 s Controllable Duration

duration accepts whole seconds from 4 to 15, or -1 to let the model pick a length (billed by actual output). Fixed 24 fps.

Multilingual Prompts

Chinese (up to ~500 chars) and English (up to ~1000 words), plus Japanese, Spanish, Portuguese, and Indonesian.

First+Last / First Frame

Pin both first and last frames with two images, or animate a single image as the first frame. Combine with return_last_frame to chain clips into longer continuous videos.

Multi-modal Reference-to-Video

Mix 0-9 reference images with 0-3 reference videos and 0-3 reference audios (at least 1 image or 1 video; the three image modes are mutually exclusive) to create, edit, or extend videos while keeping characters and style consistent.

Async Task Flow

Submit and get a task_id, poll for status, then download the mp4 from content.video_url (link valid for 24 hours).

Reproducible Seeds

Fix seed for similar results across runs. watermark defaults to false — output is watermark-free.

Pricing

Pricing in one line — precise token billing, tier-by-tier pegged to the Volcengine official site. The three models are priced differently: mini < fast < standard (same direction as the official site; mini runs at about half the standard model’s unit price — they are NOT the same price level). On-platform list price runs about 1.1× the official list, and with the top-up bonus (10% for general customers, up to 20% for large-deposit customers) the effective cost is essentially on par with the official site — at some tiers (e.g. 1080p large-customer price) even lower. Billing is by area×duration, so a ±5% deviation is normal — you’re welcome to test, reconcile, and reach out anytime.
Token-based billing: tokens ≈ (input video duration + output duration)(s) × output width × output height × 24 / 1024 (input video duration is 0 for text-/image-to-video; verified in our tests to within 0.1%). Since every ratio in a tier has the same pixel area, price depends only on the resolution tier, output duration, and whether the input includes video.

Official price anchors (16:9 / 5 s output, CNY per video)

① No input video (text-to-video / image-to-video / reference images): ② With input video (multi-modal reference incl. video_url; input video 2-15 s, low end ≈ 2-4 s input, high end ≈ 15 s input):
With input video, billed duration = input video duration + output duration, so it costs more than plain text-/image-to-video; a minimum-token floor also applies (very short inputs bill at the floor). The authoritative usage is the returned usage.completion_tokens.
Measured on-platform comparison (tested 2026-06 and 2026-07, 16:9 / default audio / no input video; CNY at the fixed 1:7 rate, for reference only):
The three models are NOT the same price — never treat them as equal. At the same resolution/duration, per-token price goes mini < fast < standard (e.g. 720p/5s: mini ≈ ¥3.16, fast ≈ ¥5.08, standard ≈ ¥6.35), matching the official price ladder. For batch production, mini saves the most money and time (our 2026-07 tests measured its effective unit price at exactly the platform’s nominal rate, 0.00% deviation); 1080p is standard-only.
Note: “CNY” is the on-platform list charge; “General ÷1.1” and “Large customer ÷1.2” are the effective prices after a 10% / 20% top-up bonus — after the bonus, prices land close to the official reference, and the 1080p large-customer price is even below official. The authoritative usage is the returned usage.completion_tokens.
Billing notes:
  • Final charges follow the console’s model pricing and call logs
  • Tasks are pre-charged on submission and settled on completion — your balance fluctuates briefly; reconcile against call logs, where one video produces two charge entries (see “Reading charges in the logs” below)
  • Rejected requests (HTTP 400 parameter errors, etc.) are not billed (verified)
  • Cost scales linearly with duration: a 15 s video costs about 3× a 5 s one

Reading charges in the logs (pre-charge + settlement)

Open the console log page at api.apiyi.com/log and search for the model name doubao-seedance-2-0 to see every charge. One video produces two charge entries:
  1. Pre-charge: an estimated amount deducted when the task is submitted (log entry labeled “non-streaming”, showing the token and group) — $0.449998 in the screenshot below
  2. Settlement (charge or refund): after the task completes, the difference is settled against the actual generated tokens (log entry labeled “streaming”, with a completion-token count) — $5.611858 below; 1080p usually incurs an additional charge
APIYI log page showing the two charge entries for one Seedance 2.0 video: pre-charge and settlement

Two charge entries for one 15 s 1080p video: pre-charge + settlement

The settlement entry shows neither the token nor its group — this is normal. The sum of the two entries is the video’s total cost.
How to read the time fields:
  1. The first entry’s (pre-charge) timestamp is the video’s submission time; its “first byte” value is how long the submission took to return a task ID (e.g. 首字节:3秒 / first byte: 3 s) — not the generation time
  2. The settlement entry shows 流式 (streaming) and 首字节:<1秒 (first byte under 1 s) — these are just internal markers on the settlement record, not a sign of any problem
  3. The video’s actual generation time is the “耗时” (elapsed) column on the “Async tasks” page (api.apiyi.com/task) in the top navigation
Reading the time and first-byte fields on the log page: the first entry is the submission time and submission latency

The first log entry's timestamp = submission time, and its first-byte value (3 s) is the submission latency; this fast example settled as a refund (negative amount), total cost 0.360000 − 0.022750 = 0.337250 USD

The Async tasks page showing each video task's submission time and generation elapsed time

The elapsed column on the Async tasks page is the actual video generation time, e.g. 158 s, 303 s

For the 15 s 1080p video in the first screenshot, the total cost = 0.449998 + 5.611858 = $6.061856. The matching task parameters are visible under “Async tasks” at the top of api.apiyi.com/task, and they line up exactly with the charges:
The 732,108 completion tokens ≈ 15 × 1248 × 1664 × 24 / 1024 (a 3:4 video at 1080p outputs 1248×1664) — consistent with the billing formula.
This 15 s 1080p video totals about ¥42.4 (nominal charge at the fixed 1:7 rate); with the top-up bonus the effective cost is roughly ¥35-39, versus an official reference of about ¥37.2 for the same spec. The official pricing itself is not cheap — cost is driven by model + resolution + duration (switching to fast / 720p / 5 s is far cheaper). This model is supplied at a thin margin to secure availability, and large-deposit customers get bigger discounts.
Beta-supply notice: Seedance 2.0 is currently in a beta supply phase. If your actual charges deviate noticeably from the table above, contact customer support and we will reconcile. Pricing will be adjusted dynamically with upstream policy (e.g. if a lower-priced official variant ships later) and APIYI’s supply capacity; capable channel partners are welcome to reach out. This model is priced to secure supply and serve customers, not for profit.

Group Setup

Seedance 2.0 runs on the dedicated SeeDance2 group (0.18x rate, CNY-denominated), with two hard requirements: ① the Token’s billing model must be Pay-as-you-go Priority (or Pay-as-you-go) — Pay-per-request tokens cannot route; ② the Token must have the SeeDance2 group enabled. Tokens on the Default group or other video groups will fail with “no available channel for this model”.
Why 0.18x? The system’s built-in unit prices for Seedance 2.0 match Volcengine’s official list prices — but that list is denominated in CNY, while APIYI balances are denominated in USD (fixed 1:7 USD/CNY rate). A 1x rate would effectively charge 7× the official number, so the group rate is lowered to absorb the currency conversion: 0.18 × 7 = 1.26, i.e. the nominal charge is about 1.26× the official CNY price. After stacking the recharge bonus, regular users pay roughly 10% above the official price, while high-tier recharge customers land at parity or even below it (e.g. the 1080p tier).Please be aware: billing is always based on actual token usage, and token conversion carries a small natural variance (±5% is normal); the official list price is only a reference anchor, not a per-request guarantee. The current pricing is a reasonable supply-first arrangement — always evaluate it together with the recharge bonus. If a charge looks off, we’re happy to reconcile bills with you anytime; however, “why is it slightly above the official price” is not up for debate — please keep this in mind, and skip this channel if that is a concern. On the flip side, ample concurrency with no queuing is exactly what this channel delivers.
Two recommended Token configurations:
For production we recommend B (dedicated Token): clean billing, per-line quota control, and easier troubleshooting when usage spikes.

Technical Specs

API Endpoints

Domains: api.apiyi.com is the primary gateway; vip.apiyi.com and other platform domains behave identically. The path prefix is /seedance/api/v3do not drop the /api segment, and do not use /v1/videos.

Resolutions & Aspect Ratios in Detail

A resolution tier defines the pixel area, not the short side. Actual output dimensions per ratio (official values, verified in our tests):

How adaptive works

  1. Text-to-video: the model infers the best ratio from your prompt
  2. First+last / first frame: matches the first-frame image’s ratio (mismatched images are center-cropped)
  3. Multi-modal reference-to-video: follows prompt intent, otherwise the first media item (video takes priority over images)
  4. The actual ratio used is returned in the task response’s ratio field
ratio only accepts the 7 enum values above — passing e.g. "2:1" returns an InvalidParameter error (verified), as does a duration outside 4-15. Neither is billed.

Best Practices

1

Pick the model by output needs

Choose the standard model doubao-seedance-2-0-260128 for 1080p or maximum quality; choose the lite model doubao-seedance-2-0-mini-260615 for batch production and cost-sensitive workloads (about half the standard price and the fastest generation, capped at 720p); choose fast as the middle ground.
2

Use adaptive to avoid cropping

For image-to-video keep the default adaptive so the model matches your source image’s ratio. Lock 9:16 (portrait) or 16:9 (landscape) only when the target platform demands it.
3

Duration is your cost dial

Cost scales linearly with length. Validate prompts with 5 s clips first, then scale to 10-15 s; use duration: -1 when pacing is best left to the model.
4

Turn audio off when you don't need it

generate_audio defaults to true. Pass false for silent footage you plan to score yourself.
5

Quote dialogue for better voice-over

Put spoken lines inside double quotes in the prompt — the model generates matching voices automatically.
6

Add Accept-Encoding: identity in HTTP clients

The gateway labels responses content-encoding: gzip while the body is uncompressed; auto-decompressing clients such as Python requests raise ContentDecodingError. Adding the Accept-Encoding: identity header avoids this (curl is unaffected).
7

Poll every 15-30 s and download immediately

Tasks typically finish in 2-5 minutes. content.video_url is a signed link valid for 24 hours — copy the file to your own storage as soon as the task succeeds.
8

Chain clips with return_last_frame

Set return_last_frame: true to get a watermark-free last-frame png, then use it as the next task’s first frame to build continuous multi-clip videos.

Error Codes & Retries

Client recommendations:
  • 30-60 s request timeouts are enough for create/poll calls (the wait happens on the task side)
  • Poll every 15-30 s with an overall budget of 15+ minutes (longer for 1080p / 15 s tasks)
  • Apply exponential backoff on 5xx and timeouts (2 retries)
  • Log the task id and the x-request-id response header for troubleshooting

FAQ

The most common Seedance 2.0 error: your Token does not have the SeeDance2 group enabled. Tokens on the Default group or other video groups cannot route to this model. Enable the SeeDance2 group in Token Settings and use the Pay-as-you-go Priority billing model.
The gateway’s content-encoding: gzip header does not match the actual body encoding. Symptoms include ContentDecodingError, a truncated non-JSON body (e.g. the leading {" is lost and you only get id":"cgt-xxx"}), or intermittent 400s. Add "Accept-Encoding": "identity" to your request headers; curl and browser fetch are unaffected.
generate_audio defaults to true (verified): the model adds voice, sound effects, and background music automatically. Pass "generate_audio": false explicitly for silent output.
On success the URL is at content.video_url in the poll response (not top-level). It is a signed link valid for ~24 hours — download and re-host it immediately. The task_id itself remains queryable for 7 days.
The state machine is queued → running → succeeded / failed / expired. The success state is succeeded, not completed — an easy mistake when migrating from other video APIs.
No. Seedance 2.0 rejects reference images/videos containing real human faces (upstream content safety). Alternatives: reuse face-containing output generated by Seedance models within the last 30 days, use the platform’s preset virtual avatars (asset:// IDs), or use licensed face assets.
Parameter rejections (HTTP 400) are not billed (verified). Billing is pre-charged on submit and settled on completion, so your balance fluctuates briefly — reconcile against call logs.
tokens ≈ duration(s) × width × height × 24 / 1024, verified to within 0.1%. Every ratio in a tier has the same pixel area (720p 16:9 and 9:16 both cost 108,900 tokens per 5 s) — landscape, portrait, and square all cost the same.
Price and speed go mini < fast < standard (720p/5s nominal on-platform: about ¥3.16 / ¥5.08 / ¥6.35). Pick mini for batch production and cost-sensitive workloads — about half the standard price and the fastest generation (measured 2026-07: ~1.5-2.5 min for 5 s @720p). Pick standard for 1080p or maximum detail, and fast as the middle ground. Both mini and fast cap at 720p — requesting 1080p returns a 400 parameter error (not billed).
The model picks a length between 4 and 15 s (our test produced a 10 s video) and bills by actual output. The final length is returned in the task’s duration field. Fix the duration explicitly if cost predictability matters.
No. frames and camera_fixed are Seedance 1.x parameters — not supported by the Seedance 2.0 series. Use whole-second duration instead.
No — they are three mutually exclusive modes: first+last (2 images with required first_frame/last_frame roles), first frame (1 image), and multi-modal reference-to-video (0-9 images + 0-3 videos + 0-3 audios, at least 1 image or 1 video, image role reference_image). To approximate “first/last frame + reference”, use reference mode and designate a frame via the prompt.
The SeeDance2 group has ample concurrency with no queuing (15 simultaneous tasks all ran immediately in our test). Contact sales for larger sustained workloads.
Keep prompts under ~500 Chinese characters or ~1000 English words — longer prompts dilute detail. Supported languages: Chinese, English, Japanese, Spanish, Portuguese, Indonesian. Describe subject + action + camera movement + lighting/style.
Seedance 2.0 is one of the few first-tier 2026 video models that outputs synchronized audio by default. Combined with same-price aspect ratios and a 15-second ceiling, it is a strong primary channel for short-video and e-commerce asset production. To compare alternatives, the same Token (with extra groups enabled) can call Sora 2, VEO 3.1, and Wan2.7 directly.