You should not need three API keys to generate one image. You need a generation API that returns a hosted HTTPS URL in the response. Extra keys show up because most image models emit bytes, not an address, and a coding agent cannot paste a local path into a README, a pull request, or a teammate's machine. One of those keys is often a file host that exists only to hold the output.
Three keys, one PNG
A published community MCP server — Image Toolkit, listed for Claude Desktop and Cursor — asks you to configure three environment variables before it will generate a picture: a Google Gemini key, a FreeImage.host key, and a Remove.bg key. Generation, hosting, cutout. Three sign-ups. Three dashboards. Three ways the tool call fails.
The README is not confused. It is accurate. Gemini produces the pixels. Remove.bg strips a background if you ask. FreeImage.host is there because the pixels have nowhere to live. The hosting key is not a creative feature. It is a bucket with an HTTP API.
Background removal is optional productization. You can generate a usable image without it. You cannot hand that image to an agent, a Markdown file, or a browser without an address. The interesting key is the hosting one, because it is doing infrastructure.
Other community servers look cleaner on the env sheet. One Replicate key wrapping Flux. One OpenAI key wrapping an image model. The count is one. You still catch a file. Some write to disk. Some return a short-lived URL. Some stuff base64 into the tool result. The products you operate are still two.
Why the file has to live somewhere
An image model does one job. It takes a prompt — maybe a reference image, an aspect ratio, a size — and it writes a raster. That raster is a few megabytes of PNG or JPEG. It is not a URL. It is not an <img> tag. It is bytes.
A coding agent lives in text. The next step after "generate" is almost never "admire a local preview." It is embed in a README, drop into a GitHub comment, load in a browser, ship to a teammate, or assert in CI that the asset exists. None of those consumers can open /tmp/out.png on the process that happened to run the MCP server.
So the pipeline is three stages even when the docs mention one:
- Call a model.
- Catch the bytes.
- Give those bytes an HTTPS URL.
Stage three is the one tutorials skip. It is also the one that grows an S3 bucket, an R2 bucket, CORS, a public ACL, cache headers, and a lifecycle rule. Or it grows a FreeImage.host key, which is the same stage with someone else's terms of service.
Agents make this worse. A human with a notebook can glance at a preview window. An agent in Cursor has to return something the chat, the editor, and the rest of the toolchain can fetch later. MCP can return binary inline. Clients are picky about mime types, and a multi-megabyte payload is a bad context citizen. A URL is the format the web already uses.
If the generator hands you a URL that expires, you have not skipped hosting. You have scheduled a 404. Copying the object to your own bucket before the link dies is still a file-hosting pipeline. It just fails in production, in a README, on a Tuesday.
What people bolt on instead
Local disk. Fast. Fine on a laptop. Broken the moment the process is a CI runner, a remote MCP host, or a colleague's clone. A path is not an API.
Base64 in the tool result. No second account. The conversation now carries the entire file. Every follow-up pays for those tokens again. You still do not have a URL you can paste.
A temporary URL from the generator. The shape looks right. Then it expires, and every consumer that bookmarked it is wrong. You add object storage to persist the object. Back to two products, plus a race against TTL.
A dedicated upload API — FreeImage.host, Imgur, S3, R2. This is the Image Toolkit pattern. It works. It also means "generate image" is two network calls, two auth headers, two rate limits, and two failure modes. If the upload fails after generation succeeded, you paid for pixels you cannot show.
A specialist cutout API on top. Remove.bg is a real product. It is not why the stack is three keys. It is why the README looks busier. Treat it as a feature you might want. Do not treat it as a requirement of producing a PNG.
There is also the ops tax. Three keys means three rotation policies, three leak hunts, three status pages, and an agent that can fail because the host is down even when the model is fine. Auth errors get blamed on the generator. The generator was fine. The bucket rejected the PUT.
If you need an air-gapped GPU or a private fine-tune, assembling this yourself is reasonable. You already run storage. For "my agent should return a picture URL," you are signing up for a host because the generator refused to.
The dependency that should not exist
The hosting key disappears when the generation API is also the file host. Not as a second product you configure. As the response.
You POST a prompt. You get JSON back. One of the fields is an HTTPS URL on a CDN the API already runs. You do not create a bucket. You do not mint a presigned PUT. You do not copy bytes off a temp link. The agent embeds the URL. The README fetches it. An image view loads it. Same string.
That is a smaller surface than it sounds. Auth is one header. Billing is one meter. Errors are one HTTP status. The tool schema the agent sees is generate-in, URL-out. A follow-up edit can take that URL as input instead of asking the user to re-upload from disk.
Synchronous image generation matters here. If a simple generate call can wait and return the URL in the same body, the agent does not poll and does not register a webhook for the common case. Video and batch jobs take longer. Async plus a webhook is the honest contract for work that will not fit in one request. Mixing those up is how you get retry loops around a job that already finished, or a chat that hangs on a render that was never going to.
You give up something. The pixels leave your machine and live on the provider's CDN. If that is unacceptable, run a local model and keep your own store — you are back to operating stage three. A hosted URL is also not a substitute for a cutout specialist if you truly need that model. Hosting is infrastructure. Background removal is a feature. Do not conflate them just because a README listed both as env vars.
One REST API, one key
PicX is built around that response shape. One REST API, base URL https://api.picxstudio.com, one key with a pxsk_ prefix, Authorization: Bearer $PICX_API_KEY. Roughly 33 image and video models sit behind that key. Billing is in credits per generation, never a currency amount on the meter, and you can set a hard spend cap on a key so an agent cannot drain the account because a prompt looped.
Image generation is synchronous: a generate call returns a hosted image URL in the response body. There is no poll on that path. We store the file on PicX's own CDN. You do not bring a bucket. You do not configure a separate host. Video and batch jobs are asynchronous and delivered by webhook.
The agent-facing surface is an MCP server with 19 tools, documented at https://picxstudio.com/developers/mcp. The same platform ships Agent Skills at /skills, a CLI, official Python and JavaScript SDKs on npm and PyPI as picx-ai, an interactive playground, live `llms.txt` and `llms-full.txt`, and `/agent-setup/prompt.md`, an executable checklist a coding agent can run. Connection pages exist for Claude Desktop, Cursor, and ChatGPT.
You keep one secret in the environment. Not three. The MCP client shows the tool call. You approve it. A URL comes back. If you would rather skip MCP, the REST API, the CLI, and the SDKs hit the same generate endpoint and get the same hosted URL.
Spend caps do not make a bad prompt cheap. They bound the blast radius of one key. That is a different problem from collapsing the host. Both matter when the caller is an agent that will happily retry.
What this does not excuse
A hosted URL does not make the language model a photographer. Bad prompts still produce bad pictures. Iterate with an edit tool that accepts the previous URL.
It also does not mean you should generate a raster when SVG is the right file. Icons, diagrams, and anything that must theme with CSS still belong in vectors.
Do not put a pxsk_ key in a public client or a committed env file. One key is easier to leak than three. Scope it, cap it, keep it on a server or in the MCP client's environment, not in an App Store binary.
If you already run object storage for user uploads, keep using it for those uploads. Generated output does not have to share that bucket. Mixing files your users sent you and files a model just made in the same lifecycle policy is how a hero image disappears on a Tuesday.
FAQ
Why do some image MCP servers require a separate file-hosting API key?
Because the generation API they wrap returns bytes or a short-lived link, not a URL the agent can paste into Markdown. A second service — often a public image host — exists only to give those bytes an HTTPS address.
Can I write the image to disk and skip hosting?
Only on your laptop — a local path fails in CI, on a teammate's machine, in a GitHub README, and in any MCP client that is not the process that wrote the file. If the next consumer is the web, you need a URL.
Does PicX require me to configure S3, R2, or another file host?
No — output files are hosted on PicX's own CDN, and a simple image generate call returns that URL in the response body. You do not bring or configure a separate file-hosting service.
Is a specialist background-removal API still a separate product?
Yes, if you want that specific cutout model — it is a feature, not a host. You should not need a third key whose only job is to store the PNG the generator already produced.
How do I stop an agent from spending across several image services?
Collapse generation and hosting onto one key first, then set a hard spend cap on that key. PicX bills in credits per generation and lets you cap a pxsk_ key so a retry loop cannot drain the account.



