# Generate Video

> Submit an asynchronous video job, then poll the generation for the result.

> [!WARNING]
> Video generation is asynchronous and can take from ~30 seconds to several minutes. Submit a job, then poll for status — or register a webhook to be notified on completion.

## `POST /v1/videos/generate`

Submit an async video generation job. Responds 202 Accepted with a generation id and a poll_url.

**Auth:** API Key

| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `prompt` | `string` | Conditional | Text description of the video. Required for every mode **except** `lipsync`, where the speech comes from `audio_url` instead. Maximum 4000 characters. |
| `model` | `string` | No | Video model id. Defaults to fal-ai/bytedance/seedance/v2. |
| `mode` | `string` | No | Generation mode: text, image, reference, frames, extend, lipsync, or edit. Default text. See the mode table below for the fields each one requires. |
| `duration` | `number` | No | Duration in seconds, 1 to 60. Default 5. Sent and priced for every mode, including lipsync. |
| `resolution` | `string` | No | Output resolution: 480p, 720p, or 1080p. Default 720p. Sent and priced for every mode, including lipsync. |
| `aspect_ratio` | `string` | No | Aspect ratio hint, e.g. `16:9`. Maximum 10 characters. |
| `sound` | `boolean` | No | Include audio. Default true. Sent and priced for every mode, including lipsync. |
| `image_url` | `string` | No | Source image URL. Required for `image` mode, and for `edit` mode (unless `reference_urls` is supplied). Maximum 2048 characters. |
| `reference_urls` | `string[]` | No | Reference image URLs — up to 10. Required for `reference` mode. |
| `start_frame_url` | `string` | No | First frame. Required for `frames` mode. Maximum 2048 characters. |
| `end_frame_url` | `string` | No | Last frame. Optional for `frames` mode. Maximum 2048 characters. |
| `source_video_url` | `string` | No | Source video. Required for `extend`, `lipsync`, and `edit` modes. Maximum 2048 characters. |
| `audio_url` | `string` | No | Audio track. Required for `lipsync` mode. Maximum 2048 characters. |
| `webhook` | `object` | No | Inline webhook binding. See webhooks section below. |
| `callback_url` | `string` | No | Delivery URL to POST the result to on completion. Maximum 2048 characters. |
| `idempotency_key` | `string` | No | Deduplicate submissions; the same key returns the same job. |

### Generation modes

`mode` selects what drives the render. Each mode has its own required inputs — a request missing them is rejected with `422` before any credits are spent.

| Mode | Required fields | What it does |
| --- | --- | --- |
| `text` | `prompt` | Generate a video from a text prompt alone. |
| `image` | `prompt`, `image_url` | Animate a single source image. |
| `reference` | `prompt`, `reference_urls` | Generate a video guided by up to 10 reference images. |
| `frames` | `prompt`, `start_frame_url` | Interpolate from a start frame; `end_frame_url` is optional. |
| `extend` | `prompt`, `source_video_url` | Continue an existing video. |
| `lipsync` | `source_video_url`, `audio_url` | Drive the lips in a source video from an audio track. `prompt` is **not** required — the speech comes from the audio. |
| `edit` | `prompt`, `source_video_url`, and `image_url` (or `reference_urls`) | Edit a source video guided by an image. |

> [!NOTE]
> `duration`, `resolution`, and `sound` are read for every mode, including `lipsync` — the job is priced from them regardless of which mode you pick.

**Response**

```json
{
  "id": "b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f",
  "status": "pending",
  "type": "video",
  "model": "fal-ai/bytedance/seedance/v2",
  "poll_url": "/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f",
  "events_url": "/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f/events",
  "webhook": {
    "mode": "none",
    "id": null,
    "url": null,
    "events": []
  }
}
```

## `GET /v1/generations/{id}`

Poll a generation by id. Status transitions: pending → processing → completed or failed.

**Auth:** API Key

**Response**

```json
{
  "id": "b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f",
  "status": "completed",
  "type": "video",
  "model": "fal-ai/bytedance/seedance/v2",
  "output_url": "https://cdn.picxstudio.com/videos/b3f1c8e2.mp4",
  "error_message": null,
  "credits_used": 120,
  "created_at": "2026-06-17T20:00:00Z",
  "completed_at": "2026-06-17T20:02:30Z"
}
```

## Submit and poll

**JavaScript SDK**

```bash
npm install picx-ai
```

```js
import { PicX } from "picx-ai";

const picx = new PicX(process.env.PICX_API_KEY);

const job = await picx.video.create({
  prompt: "a cat walking in the rain, cinematic",
  duration: 5,
  resolution: "720p",
  sound: true,
});

console.log(job.id, job.status); // "pending"

// Polls with backoff until the video is ready
const asset = await job.wait();
console.log(asset.url);
```

**Python SDK**

```bash
pip install picx-ai
```

```python
import os
from picx import PicX

picx = PicX(os.environ["PICX_API_KEY"])

job = picx.video.create(
    prompt="a cat walking in the rain, cinematic",
    duration=5,
    resolution="720p",
    sound=True,
)

print(job.id, job.status)  # "pending"

# Polls with backoff until the video is ready
asset = job.wait()
print(asset.url)
```

**curl**

```bash
# 1. Submit the async job
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"a cat walking in the rain, cinematic","duration":5,"resolution":"720p","sound":true}'

# 2. Poll the generation until it completes
curl https://api.picxstudio.com/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f \
  -H "Authorization: Bearer $PICX_API_KEY"
```

## Modes in practice

Every request below submits the same async job and returns the same `202 Accepted` body shown above; only the inputs differ.

**image** — animate a single source image:

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "image",
    "prompt": "slow zoom out, gentle wind",
    "image_url": "https://cdn.picxstudio.com/uploads/api/scene.png",
    "duration": 5,
    "resolution": "720p"
  }'
```

**reference** — guide the render with up to 10 reference images:

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "reference",
    "prompt": "the same character walking through a market",
    "reference_urls": [
      "https://cdn.picxstudio.com/uploads/api/ref1.png",
      "https://cdn.picxstudio.com/uploads/api/ref2.png"
    ],
    "duration": 5
  }'
```

**frames** — interpolate from a start frame (end frame optional):

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "frames",
    "prompt": "smooth transition between the two frames",
    "start_frame_url": "https://cdn.picxstudio.com/uploads/api/first.png",
    "end_frame_url": "https://cdn.picxstudio.com/uploads/api/last.png",
    "duration": 5
  }'
```

**extend** — continue an existing video:

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "extend",
    "prompt": "the car keeps driving down the coast road",
    "source_video_url": "https://cdn.picxstudio.com/uploads/api/clip.mp4",
    "duration": 5
  }'
```

**lipsync** — drive the lips in a source video from an audio track. No `prompt` — the speech comes from `audio_url`. `duration`, `resolution`, and `sound` are still sent, because the job is priced from them:

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "lipsync",
    "source_video_url": "https://cdn.picxstudio.com/uploads/api/speaker.mp4",
    "audio_url": "https://cdn.picxstudio.com/uploads/api/voiceover.mp3",
    "duration": 5,
    "resolution": "720p",
    "sound": true
  }'
```

**edit** — edit a source video guided by an image:

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mode": "edit",
    "prompt": "restyle the scene as a watercolor painting",
    "source_video_url": "https://cdn.picxstudio.com/uploads/api/clip.mp4",
    "image_url": "https://cdn.picxstudio.com/uploads/api/style.png",
    "duration": 5
  }'
```

> [!TIP]
> The `image`, `reference`, `frames`, `extend`, `lipsync`, and `edit` modes all take **public** URLs. If a source file only lives on your device, upload it through [Managed Assets](/docs/api-reference/managed-assets) first and use the returned `url`.

## Poll manually

If you store the job ID and check it from a queue worker or a webhook handler:

```js
const job = await picx.video.create({ prompt: "a cat walking in the rain" });

// Persist job.id, then later, from anywhere:
let generation = await picx.generations.get(job.id);

while (generation.status !== "completed" && generation.status !== "failed") {
  await new Promise((r) => setTimeout(r, 5000));
  generation = await picx.generations.get(job.id);
}

if (generation.status === "failed") throw new Error(generation.error);
console.log(generation.url);
```

## Stream events via SSE

The `events_url` in the response supports Server-Sent Events for real-time status updates without polling:

```bash
curl -N https://api.picxstudio.com/v1/generations/b3f1c8e2-1a2b-4c3d-9e8f-0a1b2c3d4e5f/events \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Accept: text/event-stream"
```

Each SSE frame is a JSON snapshot of the generation state. The stream closes when the job reaches a terminal status (completed, failed, or cancelled).

## Inline webhook binding

Bind a webhook directly in the generation request so you are notified on completion without polling:

```bash
curl -X POST https://api.picxstudio.com/v1/videos/generate \
  -H "Authorization: Bearer $PICX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "a cat jumping off a table",
    "duration": 5,
    "resolution": "480p",
    "webhook": { "url": "https://your-server.com/webhook" }
  }'
```

```js
const job = await picx.video.create({
  prompt: "a cat jumping off a table",
  duration: 5,
  resolution: "480p",
  webhook: { url: "https://your-server.com/webhook" },
});
// job.raw.webhook.mode === "inline"
```

```python
job = picx.video.create(
    prompt="a cat jumping off a table",
    duration=5,
    resolution="480p",
    webhook_url="https://your-server.com/webhook",
)
# job.raw["webhook"]["mode"] == "inline"
```

> [!NOTE]
> For production, prefer a registered webhook or an inline webhook binding over polling — you are notified the moment the job finishes. See Developer Tools → Webhooks.

## FAQ

### Is video generation synchronous like image generation?

No — video is asynchronous. `POST /v1/videos/generate` responds `202 Accepted` with a job id right away; you then poll `GET /v1/generations/{id}` (or stream SSE, or use a webhook) until the status reaches `completed` or `failed`. It can take from ~30 seconds to several minutes.

### How do I get notified without polling in a loop?

Either stream `events_url` via Server-Sent Events for real-time status, or bind a `webhook` in the generation request (or register one — see [Webhooks](/docs/developer-tools/webhooks)) and get a POST the moment the job finishes.

### What video durations and resolutions are supported?

Duration is 1–60 seconds (default 5). Resolution is 480p, 720p, or 1080p (default 720p). Sound is included by default and can be turned off with `sound: false`. All three are sent and priced for every mode, including `lipsync`.

### Which modes are available and what do they need?

Seven modes: `text` (prompt only), `image` (prompt + `image_url`), `reference` (prompt + `reference_urls`, up to 10), `frames` (prompt + `start_frame_url`, optional `end_frame_url`), `extend` (prompt + `source_video_url`), `lipsync` (`source_video_url` + `audio_url`, no prompt), and `edit` (prompt + `source_video_url` + `image_url` or `reference_urls`). A request missing a mode's required fields is rejected with `422` before any credits are spent.

### Why is `prompt` rejected as required when I use lipsync?

It isn't — `lipsync` is the one mode where `prompt` is optional, because the speech is taken from `audio_url`. Every other mode requires a non-empty `prompt`. If you send a `text`/`image`/`reference`/`frames`/`extend`/`edit` request without a prompt you get a `422`.

### Can I generate a video from a reference image instead of just a text prompt?

Yes — set `mode` to `image` and pass `image_url` for a single source image, or `mode` to `reference` and pass `reference_urls` (up to 10) to guide the render with several references.

### How do I avoid submitting the same video job twice?

Pass an `idempotency_key` in the request body — a resubmission with the same key returns the same job instead of starting a duplicate.
