You have a product shot due Thursday. Cotton henley, window light, no styling department. The knit has to read as cotton, not as a plastic suggestion of fabric. You open the model picker and get names. Seedream. Nano Banana. Flux. GPT Image 2. Ideogram. Recraft. Video is a second list: MiniMax H3, H3 Max, Seedance, Kling, Wan, Grok Imagine. None of them is labelled with the job it was built to do.
That is the actual problem. Not a shortage of models. A shortage of matching.
PicX Studio runs 33 models behind one interface, billed from a single credit balance. Switching is free in subscription terms. The cost of a wrong pick is the credits spent on a result you cannot use. Technically fine. Stylistically wrong. That waste is quieter than a failed generation, and easier to repeat.
A longer list does not make a better image
A long dropdown feels like competence. It is not. It is inventory. Each name is a different team, a different training set, a different thing they were willing to sacrifice. Photorealism costs you something. Batch consistency costs you something else. Native audio on a clip is not a bonus on every family. It is a design choice that changes what the clip is.
People treat the list as a quality ranking. First name, familiar name, last thing they saw recommended. Prompting harder does not convert a stylised model into a studio photograph. It converts a wrong tool into a more expensive wrong tool.
Name what the output has to survive. A catalog page where cotton has to look like cotton. Ten banners that have to look like one campaign. A clip where a face cannot drift. Then pick. The model is downstream of the job.
Matching is the skill, not mastery of one model
We default to a favourite. Everyone does. You learn a model's habits, you stop losing time on the picker, you tell yourself you have a look. That works until the job changes and the look does not.
A model you have mastered for illustrated campaign frames will still give you illustrated campaign frames when the brief is a believable headshot. The failure is not lack of practice. It is using practice as a substitute for selection.
The work is classification. Still or motion. Photograph or graphic. One image or a set. Sound, or a silent plate you will score later. Once those are named, the list shrinks. That shrinkage is the point.
You will still be wrong sometimes. Models overlap. A vague prompt will waste a good model as readily as a bad one. It depends, and it depends on the constraint you cannot relax. Name that first. Then generate.
Still images split on photograph, set, and structure
The still-image families are not flavours of the same generator. Their priorities show up in the jobs that fail.
Photorealism is a specific demand. Skin that looks like skin. A shadow that softens at the edge because light does that. Fabric that drapes instead of hovering. Google's Nano Banana and Gemini image models, and OpenAI's GPT Image 2, are the names we reach for when the output has to pass as a photograph. That is a working rule, not a score. The henley on a catalog page is this job. If the knit reads as illustration, you did not get a draft. You got the wrong category of object.
Batch consistency is a different demand. Ten banners. Same palette, same kind of composition, same product, different copy or crop. Looking like one campaign matters more than any single frame looking precious. ByteDance's Seedream family is where we start for that kind of set. The failure mode of a campaign is drift: one image glossy, the next painterly, the third with a different white point. A beautiful outlier is still a break in the set. Ideogram and Recraft are in the same list. They do not change the split.
Structural accuracy is the third demand, and people skip it because it is less glamorous. A glass on a table has to cast a shadow that matches the window. Reflections have to respect the room. Wood grain has to sit on the wood, not float as a sticker. Flux, from Black Forest Labs, is the open-weight family in that still-image list. Open-weight is a trade. You give up some of the out-of-the-box polish a closed model will hand you. You gain control: spatial sense, material behaviour, a model you can steer rather than only prompt. If the job is a pretty thumbnail, that trade is annoying. If the job is an object that has to obey a room, it is the reason to pick it.
Do not flatten this into a ranking. A photoreal model can still break a campaign set. A consistent stylist can still fail a cotton close-up. Pick for the constraint you cannot edit out later.
Video splits on motion, length, sound, and identity
Video families split on different axes. A stills favourite is a poor default for a clip.
Motion realism is whether a head turn looks like a neck did it. Water that ripples instead of smearing. Cloth that shifts when a shoulder does. If the clip is a person in a room, this often decides whether you can use the take. If it is a graphic lockup with a slow push-in, you can live with less.
Clip length is a separate constraint. Some jobs are a loop. Some need a beat, a pause, a second beat. You cannot prompt a short-clip family into a scene it was not built to hold. Match the length you actually need.
Native audio is not a nice extra on every model. MiniMax H3 Max renders at 768P natively and always generates audio. Always. You do not toggle it off. If you needed a silent plate to score later, that is not a setting you forgot. That is the wrong family. If you needed sound married to the picture, it is a reason to pick it rather than a surprise. MiniMax H3, Seedance, Kling, Wan, and Grok Imagine fill out the video list. They are not backups of each other.
Character consistency is the failure we keep hitting when we ignore it. A face that is almost the same person later in the clip is not almost usable. It is a reshoot. If the job is a product orbit, you can ignore this. If the job is a person walking through a space, you cannot.
Motion, length, sound, and identity pull in different directions. That is irritating. It is also the job.
Pretty is not the same as usable
The specific failure looks like this. You needed a photograph. You reached for a stylised model because the previews were striking. You got a finished illustration. As a reviewer who tested seven models inside one editor found, the gaps between models were not a matter of finish, and a stylised pick on a photographic brief produced drawings that looked considered and still could not stand in for a photograph.
That result is easy to keep. It looks like work. It is the wrong object. A client who asked for a believable product image cannot be talked into an illustrated one because the lighting is tasteful. You already spent the credits. Switching models would have cost nothing extra. Regenerating on the right family would have. The discarded result is the actual bill.
Open-weight is a different mismatch. Flux will not always look as finished as a closed image model. If you needed polish and you picked control, you will keep refining and call the model bad. If you needed control and you picked polish, you will accept a pretty spatial lie. Both are matching errors.
Before you spend the credits
Write the job in one sentence that names the constraint you cannot relax. Not the vibe. The constraint. Cotton has to read as cotton on a catalog page. Ten banners, one campaign. A clip with a face that cannot drift. Silent plate, scored later. Audio has to arrive with the picture.
Then ask four things, in order.
Still or motion. If you are unsure, it is still. Video is for a change you cannot imply with a crop.
If it is still, is it a photograph, a campaign set, or a structural problem. Photograph: the families pointed at photorealism. Set: the families pointed at batch consistency. Structure, materials, a room that has to obey light: the open-weight still family, and accept the polish trade.
If it is motion, name the non-negotiable. Realistic movement, length, sound, or identity. If sound must not be there, do not pick a model that always generates audio. If identity is the job, do not pick on the strength of a single pretty frame.
Generate once on that pick. If the result is the wrong category of object, change family. Do not write a longer prompt at the same model. A longer prompt is for a right family that missed a detail. It is not a way to turn an illustration into a photograph, or a silent need into a model that always comes with sound.
If the result is the right category and still not good enough, then you iterate. Same family. Tighter brief. That is craft. The picker was the decision. The prompt is the refinement.
Keep a short note of what you picked and why. One line. The next time the garment is different and the list looks the same, you will not have to pretend the names were ever the skill.



