Skip to main content
POST
Image-to-video: submit a video generation task from a reference image
The interactive Playground on the right supports live debugging. Set your API Key in the Authorization field (format: Bearer sk-xxx), upload a reference image, enter a prompt, choose model / size / seconds, and send.
Scope: This page covers “generate video from a reference image” — upload one image as the starting frame / visual anchor to animate static visuals. If you don’t need a reference image, use the Text-to-Video endpoint (same path, JSON body).
⚠️ Reference image dimensions must exactly match size
  • The uploaded image’s pixel dimensions must equal the size field (e.g. size=1280x720 requires a 1280×720 image)
  • Mismatch returns 400: Inpaint image must match the requested width and height
  • Pre-crop with ffmpeg / Pillow before upload
Other notes:
  • Content-Type must be multipart/form-data (not JSON)
  • Only one file is supported; the field name is fixed as input_reference
  • Accepted formats: image/jpeg / image/png / image/webp

Code Samples

Python (OpenAI SDK Drop-In)

Python (Raw requests + multipart)

cURL

Node.js (fetch + FormData)

Browser JavaScript

Parameters Quick Reference

Detailed parameter constraints, allowed values, and examples are visible in the right-hand Playground. input_reference must be uploaded via multipart — URLs and base64 are not accepted.

Reference Image Preparation

1

Pick the target resolution

Choose size first based on your use case: portrait 720x1280, landscape 1280x720, Pro 1080p landscape 1920x1080, etc.
2

Crop locally to exact pixels

Use Pillow / ffmpeg to crop the image to the target dimensions:
Or one-line ffmpeg:
3

Pick the right format

Prefer PNG (lossless, ideal for illustrations / screenshots), JPEG for photos to save bytes, WebP if you need transparency.
4

Focus the prompt on "motion" not "appearance"

The reference image already defines the visuals. The prompt should focus on how it should animate: camera push/pull, object motion, lighting changes, character expressions, etc. Example: "Camera slowly pushes in, leaves gently swaying, sunlight flickering through branches".

Response Format

The response shape is identical to Text-to-Video: submit returns id + status: "queued", polling reports progress, completion downloads via /v1/videos/{id}/content as MP4.
⚠️ Common 400 errors
  • Inpaint image must match the requested width and height — reference image dimensions don’t match size. Most common. Validate dimensions client-side before upload
  • Invalid file format — uploaded file is not jpeg / png / webp, or is corrupted
  • Missing required parameter: input_reference — multipart field name is wrong (must be input_reference, not image or reference)
  • seconds must be one of "4", "8", "12" — passed integer 4 instead of string "4"
Image-to-video and text-to-video have the same per-second pricing (billed by seconds); uploading a reference image does not cost extra. See the pricing table.

Authorizations

Authorization
string
header
required

API Key from the APIYI console (must use Sora2官转 group + usage-based billing)

Body

multipart/form-data
model
enum<string>
default:sora-2
required

Model ID. sora-2 supports 720p only; sora-2-pro supports 720p / 1024p / 1080p

Available options:
sora-2,
sora-2-pro
prompt
string
required

Video generation prompt. Focus on how the image should animate: camera motion, object motion, lighting changes

Example:

"Animate this scene: gentle waves lapping, leaves swaying, cinematic camera push-in"

input_reference
file
required

Reference image file used as the video's starting frame / visual anchor.

  • Accepted formats: image/jpeg / image/png / image/webp
  • Dimensions must equal size, otherwise you get Inpaint image must match the requested width and height
  • Only one file is supported; field name is fixed as input_reference
seconds
enum<string>
default:4

Video duration as string enum: "4" / "8" / "12"

Available options:
4,
8,
12
size
enum<string>
default:720x1280

Output resolution. Must exactly match the input_reference image dimensions:

  • sora-2 (720p only): 720x1280 / 1280x720
  • sora-2-pro additionally: 1024x1792 / 1792x1024 / 1080x1920 / 1920x1080
Available options:
720x1280,
1280x720,
1024x1792,
1792x1024,
1080x1920,
1920x1080

Response

Task submitted, returns video_id with queued status

id
string

Task ID for subsequent polling and download

Example:

"video_abc123def456"

object
string

Object type, fixed video

Example:

"video"

model
string

Model ID used for this task

Example:

"sora-2"

status
enum<string>

Task status:

  • queued — submitted, waiting in queue
  • in_progress — generating
  • completed — done, ready to download (/v1/videos/{id}/content)
  • failed — failed (not billed), safe to retry
Available options:
queued,
in_progress,
completed,
failed
Example:

"queued"

progress
integer

Generation progress percentage (0–100), not strictly linear

Example:

0

created_at
integer

Task creation Unix timestamp (seconds)

Example:

1712697600

completed_at
integer

Task completion Unix timestamp (seconds), present only on completed status

Example:

1712697900

size
string

Actual output resolution (matches the requested size)

Example:

"1280x720"

seconds
string

Actual duration generated (matches the requested seconds)

Example:

"8"

quality
string

Quality tier (standard for sora-2, high for sora-2-pro)

Example:

"standard"