Generate Video
Submit an asynchronous video job, then poll the generation for the result.
Video generation is asynchronous and can take from ~30 seconds to several minutes. Submit a job, then poll for status — or register a webhook to be notified on completion.
POST /v1/videos/generate
Submit an async video generation job. Responds 202 Accepted with a generation id and a poll_url.
Auth: API Key
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt |
string |
Conditional | Text description of the video. Required for every mode except lipsync, where the speech comes from audio_url instead. Maximum 4000 characters. |
model |
string |
No | Video model id. Defaults to fal-ai/bytedance/seedance/v2. |
mode |
string |
No | Generation mode: text, image, reference, frames, extend, lipsync, or edit. Default text. See the mode table below for the fields each one requires. |
duration |
number |
No | Duration in seconds, 1 to 60. Default 5. Sent and priced for every mode, including lipsync. |
resolution |
string |
No | Output resolution: 480p, 720p, or 1080p. Default 720p. Sent and priced for every mode, including lipsync. |
aspect_ratio |
string |
No | Aspect ratio hint, e.g. 16:9. Maximum 10 characters. |
sound |
boolean |
No | Include audio. Default true. Sent and priced for every mode, including lipsync. |
image_url |
string |
No | Source image URL. Required for image mode, and for edit mode (unless reference_urls is supplied). Maximum 2048 characters. |
reference_urls |
string[] |
No | Reference image URLs — up to 10. Required for reference mode. |
start_frame_url |
string |
No | First frame. Required for frames mode. Maximum 2048 characters. |
end_frame_url |
string |
No | Last frame. Optional for frames mode. Maximum 2048 characters. |
source_video_url |
string |
No | Source video. Required for extend, lipsync, and edit modes. Maximum 2048 characters. |
audio_url |
string |
No | Audio track. Required for lipsync mode. Maximum 2048 characters. |
webhook |
object |
No | Inline webhook binding. See webhooks section below. |
callback_url |
string |
No | Delivery URL to POST the result to on completion. Maximum 2048 characters. |
idempotency_key |
string |
No | Deduplicate submissions; the same key returns the same job. |
Generation modes
mode selects what drives the render. Each mode has its own required inputs — a request missing them is rejected with 422 before any credits are spent.
| Mode | Required fields | What it does |
|---|---|---|
text |
prompt |
Generate a video from a text prompt alone. |
image |
prompt, image_url |
Animate a single source image. |
reference |
prompt, reference_urls |
Generate a video guided by up to 10 reference images. |
frames |
prompt, start_frame_url |
Interpolate from a start frame; end_frame_url is optional. |
extend |
prompt, source_video_url |
Continue an existing video. |
lipsync |
source_video_url, audio_url |
Drive the lips in a source video from an audio track. prompt is not required — the speech comes from the audio. |
edit |
prompt, source_video_url, and image_url (or reference_urls) |
Edit a source video guided by an image. |
duration,resolution, andsoundare read for every mode, includinglipsync— the job is priced from them regardless of which mode you pick.
Response
{
"id": "b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f",
"status": "pending",
"type": "video",
"model": "fal-ai/bytedance/seedance/v2",
"poll_url": "/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f",
"events_url": "/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f/events",
"webhook": {
"mode": "none",
"id": null,
"url": null,
"events": []
}
}
GET /v1/generations/{id}
Poll a generation by id. Status transitions: pending → processing → completed or failed.
Auth: API Key
Response
{
"id": "b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f",
"status": "completed",
"type": "video",
"model": "fal-ai/bytedance/seedance/v2",
"output_url": "https://cdn.picxstudio.com/videos/b3f1c8e2.mp4",
"error_message": null,
"credits_used": 120,
"created_at": "2026-06-17T20:00:00Z",
"completed_at": "2026-06-17T20:02:30Z"
}
Submit and poll
JavaScript SDK
npm install picx-ai
import { PicX } from "picx-ai";
const picx = new PicX(process.env.PICX_API_KEY);
const job = await picx.video.create({
prompt: "a cat walking in the rain, cinematic",
duration: 5,
resolution: "720p",
sound: true,
});
console.log(job.id, job.status); // "pending"
// Polls with backoff until the video is ready
const asset = await job.wait();
console.log(asset.url);
Python SDK
pip install picx-ai
import os
from picx import PicX
picx = PicX(os.environ["PICX_API_KEY"])
job = picx.video.create(
prompt="a cat walking in the rain, cinematic",
duration=5,
resolution="720p",
sound=True,
)
print(job.id, job.status) # "pending"
# Polls with backoff until the video is ready
asset = job.wait()
print(asset.url)
curl
# 1. Submit the async job
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"a cat walking in the rain, cinematic","duration":5,"resolution":"720p","sound":true}'
# 2. Poll the generation until it completes
curl https://api.picxstudio.com/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f \
-H "Authorization: Bearer $PICX_API_KEY"
Modes in practice
Every request below submits the same async job and returns the same 202 Accepted body shown above; only the inputs differ.
image — animate a single source image:
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "image",
"prompt": "slow zoom out, gentle wind",
"image_url": "https://cdn.picxstudio.com/uploads/api/scene.png",
"duration": 5,
"resolution": "720p"
}'
reference — guide the render with up to 10 reference images:
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "reference",
"prompt": "the same character walking through a market",
"reference_urls": [
"https://cdn.picxstudio.com/uploads/api/ref1.png",
"https://cdn.picxstudio.com/uploads/api/ref2.png"
],
"duration": 5
}'
frames — interpolate from a start frame (end frame optional):
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "frames",
"prompt": "smooth transition between the two frames",
"start_frame_url": "https://cdn.picxstudio.com/uploads/api/first.png",
"end_frame_url": "https://cdn.picxstudio.com/uploads/api/last.png",
"duration": 5
}'
extend — continue an existing video:
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "extend",
"prompt": "the car keeps driving down the coast road",
"source_video_url": "https://cdn.picxstudio.com/uploads/api/clip.mp4",
"duration": 5
}'
lipsync — drive the lips in a source video from an audio track. No prompt — the speech comes from audio_url. duration, resolution, and sound are still sent, because the job is priced from them:
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "lipsync",
"source_video_url": "https://cdn.picxstudio.com/uploads/api/speaker.mp4",
"audio_url": "https://cdn.picxstudio.com/uploads/api/voiceover.mp3",
"duration": 5,
"resolution": "720p",
"sound": true
}'
edit — edit a source video guided by an image:
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "edit",
"prompt": "restyle the scene as a watercolor painting",
"source_video_url": "https://cdn.picxstudio.com/uploads/api/clip.mp4",
"image_url": "https://cdn.picxstudio.com/uploads/api/style.png",
"duration": 5
}'
The
image,reference,frames,extend,lipsync, andeditmodes all take public URLs. If a source file only lives on your device, upload it through Managed Assets first and use the returnedurl.
Poll manually
If you store the job ID and check it from a queue worker or a webhook handler:
const job = await picx.video.create({ prompt: "a cat walking in the rain" });
// Persist job.id, then later, from anywhere:
let generation = await picx.generations.get(job.id);
while (generation.status !== "completed" && generation.status !== "failed") {
await new Promise((r) => setTimeout(r, 5000));
generation = await picx.generations.get(job.id);
}
if (generation.status === "failed") throw new Error(generation.error);
console.log(generation.url);
Stream events via SSE
The events_url in the response supports Server-Sent Events for real-time status updates without polling:
curl -N https://api.picxstudio.com/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f/events \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Accept: text/event-stream"
Each SSE frame is a JSON snapshot of the generation state. The stream closes when the job reaches a terminal status (completed, failed, or cancelled).
Inline webhook binding
Bind a webhook directly in the generation request so you are notified on completion without polling:
curl -X POST https://api.picxstudio.com/v1/videos/generate \
-H "Authorization: Bearer $PICX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "a cat jumping off a table",
"duration": 5,
"resolution": "480p",
"webhook": { "url": "https://your-server.com/webhook" }
}'
const job = await picx.video.create({
prompt: "a cat jumping off a table",
duration: 5,
resolution: "480p",
webhook: { url: "https://your-server.com/webhook" },
});
// job.raw.webhook.mode === "inline"
job = picx.video.create(
prompt="a cat jumping off a table",
duration=5,
resolution="480p",
webhook_url="https://your-server.com/webhook",
)
# job.raw["webhook"]["mode"] == "inline"
For production, prefer a registered webhook or an inline webhook binding over polling — you are notified the moment the job finishes. See Developer Tools → Webhooks.
FAQ
Is video generation synchronous like image generation?
No — video is asynchronous. POST /v1/videos/generate responds 202 Accepted with a job id right away; you then poll GET /v1/generations/{id} (or stream SSE, or use a webhook) until the status reaches completed or failed. It can take from ~30 seconds to several minutes.
How do I get notified without polling in a loop?
Either stream events_url via Server-Sent Events for real-time status, or bind a webhook in the generation request (or register one — see Webhooks) and get a POST the moment the job finishes.
What video durations and resolutions are supported?
Duration is 1–60 seconds (default 5). Resolution is 480p, 720p, or 1080p (default 720p). Sound is included by default and can be turned off with sound: false. All three are sent and priced for every mode, including lipsync.
Which modes are available and what do they need?
Seven modes: text (prompt only), image (prompt + image_url), reference (prompt + reference_urls, up to 10), frames (prompt + start_frame_url, optional end_frame_url), extend (prompt + source_video_url), lipsync (source_video_url + audio_url, no prompt), and edit (prompt + source_video_url + image_url or reference_urls). A request missing a mode's required fields is rejected with 422 before any credits are spent.
Why is prompt rejected as required when I use lipsync?
It isn't — lipsync is the one mode where prompt is optional, because the speech is taken from audio_url. Every other mode requires a non-empty prompt. If you send a text/image/reference/frames/extend/edit request without a prompt you get a 422.
Can I generate a video from a reference image instead of just a text prompt?
Yes — set mode to image and pass image_url for a single source image, or mode to reference and pass reference_urls (up to 10) to guide the render with several references.
How do I avoid submitting the same video job twice?
Pass an idempotency_key in the request body — a resubmission with the same key returns the same job instead of starting a duplicate.