Your image generation call should return the file in the same response when the job is a single still that can finish on the request. It should return a receipt you deliver later — preferably by webhook, not a poll loop — when the work is video, a batch, or any render that outlives a reasonable HTTP timeout. Those are two contracts. A well-designed API offers both. Treating every generation as a queued job makes a still cost extra state, extra calls, and extra chances for a coding agent to double-submit. Treating every generation as synchronous hangs the client on work that was never going to finish in-band.
A job id is a promise, not a file
Most generation APIs answer a POST with an identifier. You store it. You poll. You wait for a terminal status. Then you fetch a URL, or bytes, or a signed link that expires. The pattern is familiar because GPU work is variable. A queue protects the provider from holding connections open, and from your client's idle timeout.
That is a reasonable default for a video encoder. It is a weaker default for a still.
A job id is not the image. It is a promise that an image might exist later. Every consumer of that promise now owns a waiter: a loop, a sleep, a status enum, a timeout, a retry, a second code path for failure. Scripts grow a while. Backends grow a worker. Coding agents grow a tool-call loop that spends tokens re-asking whether the file exists.
The industry copies this shape from training jobs and from video. Image generate inherited it. Inheritance is not design.
Return the still in the same response
A still image is often a short, single-shot job. One prompt. One model. One file. The caller is a script, a CLI, a backend route, or an MCP tool sitting in a chat turn. The next thing that caller wants is an HTTPS URL it can embed, download, or pass to an edit.
Synchronous here means: the generate call blocks until the image exists, then returns that URL in the response body. No poll. No second endpoint. No job table on your side for the happy path.
PicX image generation works this way. One REST API at https://api.picxstudio.com. One key with a pxsk_ prefix. Auth is a single header:
Authorization: Bearer $PICX_API_KEYA simple generate call returns a hosted image URL. The file already lives on PicX's own CDN. You do not bring a bucket. You do not configure a second host. The URL in the body is the asset.
That contract matches how people actually call a still. A CLI should print a URL and exit. A backend needs the image before it writes a row. An agent tool must return something the next turn can fetch. A playground click should show a picture, not a spinner over a ticket.
Official Python and JavaScript SDKs (picx-ai on npm and PyPI) keep that path as the default for images. You call generate. You read the URL. The MCP server at https://picxstudio.com/developers/mcp is the same idea in tool form: the result that matters is still a URL, not a job you have to babysit.
Synchronous is not a claim about GPU speed. It is a shorter protocol. The model still runs. The difference is whether your process is allowed to wait on the line for that run, and whether the API is willing to sit with you.
Holding the connection has a cost
The client occupies a socket, a worker, a serverless invocation, or an editor tool-call slot for the whole render. Proxies drop idle connections. Some function runtimes kill the request before a large still finishes. If the connection dies, you can lose the result even though the work ran — unless the API also persisted the generation and you have another way to read it.
Synchronous also serializes the caller. Fan-out across many images is awkward if every call is a long POST. A laptop script does not care. A function runtime with a short idle timeout does.
There is a cost on the provider side too. Open connections are memory and load-balancer state. A fleet that only offers in-band responses will eventually lie: it will time you out, or it will pretend a long job is short. Lying is worse than a job id.
So synchronous is right when three things are true. The work is one still. The caller's runtime can wait. The failure you care about is "the response never came," not "I needed to submit a pile of jobs and go away." If any of those is false, queue it.
Queue the work that outlives the request
Video is not a still. A clip takes long enough that holding the HTTP request is malpractice. Batch is not a still either: you want to submit many jobs, persist many ids, and collect files as they land. Long renders — heavy models, anything you would not stare at a spinner for — belong on the same side of the line.
Queued means the submit call returns quickly with a receipt. The work runs elsewhere. The file shows up later.
PicX treats video and batch that way. They are asynchronous. Delivery is by webhook. That is the honest split. Pretending a video is synchronous hangs Claude Desktop, Cursor, a CI job, or your own worker until a proxy cuts the line. Pretending a still is a job makes the agent write a waiter it does not need.
What you give up: the one-call happy path. You now have two moments — submit and complete — and you must correlate them. You need somewhere to receive the completion, or you will poll. You need idempotency on submit, because a timeout on the receipt is not a timeout on the work, and a retry will otherwise start a second billable job.
Credits make that last point concrete. PicX bills in credits per generation, never a currency amount on the meter. A double-submit is two deductions, not a confused invoice line. Hard spend caps on a key bound the damage. They do not un-duplicate a job you posted twice because the first response never arrived. Persist the receipt. Treat a success with a URL as done. Do not POST generate again unless you intend to pay for a second file.
Polling is a protocol you will write badly
A queue without a delivery mechanism is half an API. You will invent the other half in every client.
Polling is the half people invent first. Store the id. Sleep. GET status. Repeat. It works from a laptop, from CI, from a worker that already has a loop. It needs no public URL. It is also waste: most of those GETs say not yet. You choose an interval. Too fast and you hammer. Too slow and you learn late. Agents are bad at this. They either poll every turn or they forget and generate again.
A webhook inverts the wait. Your server receives a POST when the job finishes. You persist the id at submit so you can match the event. You verify a signature on the raw body. You treat delivery as at-least-once, because a retry after you already handled the event is normal.
Webhooks have their own costs. You need a public HTTPS endpoint. Localhost will not do. You need to answer quickly and move heavy work off the request, or the sender will retry you. If your endpoint is down, the file can still exist — the generation succeeded — and you must have a way to read it by id. Never treat the webhook as the only copy of the result.
Polling is the right fallback and the right tool for a one-off script that should not stand up an endpoint. Webhooks are the right production path for video, batch, and any volume where a wait-loop is a product. Doing both — webhook plus a sweep for receipts that never landed — is the pattern that survives a deploy in the middle of a render.
Do not poll a still that already returned a URL. The URL is the result. A second generate "just to be sure" is a second charge.
One API, two return shapes
The design is not sync versus queue, pick a religion. The design is: pick per job class, on one API, with one auth story.
Same base URL. Same Bearer key. Same billing unit. Different return shapes, selected by the kind of work. If the API is careful, a still that cannot hold the connection can opt into the queued path with an explicit delivery target. Video stays queued. You do not make the caller learn two products, two keys, or two dashboards to get both shapes.
PicX is built on that split. Roughly 33 image and video models sit behind one pxsk_ key. Images can return in-band: generate, get a CDN URL. Video and batch go out of band: webhook delivery. Output is hosted either way. The agent-facing surfaces all talk to that one API: an MCP server (19 tools), Agent Skills at https://picxstudio.com/skills, the CLI, picx-ai, the interactive playground, live `llms.txt` and `llms-full.txt`, and `/agent-setup/prompt.md`, an executable setup checklist a coding agent can run. Connection pages exist for Claude Desktop, Cursor, and ChatGPT (https://picxstudio.com/developers/mcp).
A native iOS client is a useful test of the contract. A native app wants a URL it can load into an image view, not a job id to drive on a spinner. A video job wants a progress state and a callback into your backend, not a frozen UI thread. Same key. Different return.
Hosted output is the other half of "the response is the file." Without it, even a synchronous call still leaves you holding bytes you have to park. You give up private object storage you operate. If the pixels cannot leave your network, this shape is wrong for you: you will run a model and a bucket, and you will write the waiter yourself.
FAQ
Should my image generate call return a job id I have to poll?
Not for a simple still. Return the hosted image URL in the same response so a script, a backend, or a coding agent can use the file without a wait loop. Queue the call when the work is video, a batch, or a render that will outlive the HTTP request.
When is a webhook better than polling a generation?
When you have a public endpoint and the job is long enough that a wait-loop is a product: video, batch, many in-flight renders. Polling is fine for a one-off script with no callback URL. In production, deliver by webhook and keep a read-by-id fallback for the events you miss.
Does PicX use the same contract for images and video?
No. A simple image generate is synchronous and returns a hosted CDN URL in the response body. Video and batch are asynchronous and delivered by webhook. One REST API at https://api.picxstudio.com and one pxsk_ key cover both.
What does the generate response contain if I am not polling?
A URL on PicX's own CDN. You do not configure a separate file host. The URL is the asset: embed it, fetch it, or pass it to a later edit.


