You upload a portrait, ask for a new setting or a different lighting scheme, and the result looks expensive. The skin is smoother. The eyes are larger. The jaw is cleaner. Then you look again and realize the person is not you. The model did not fail to generate a portrait. It succeeded at generating a better-looking stranger, because your prompt invited it.
This is identity drift. Image-to-image portrait models are trained to produce faces that score well against attractive, well-lit, commercially usable images. When your instructions are vague about the person and vivid about the look, the model treats the face as the part that can be improved. Likeness is a constraint. Beauty is a reward. If you do not lock the constraint, the reward wins.
What identity drift actually looks like
Identity drift is not always a full replacement. The first pass often keeps the hair color and a rough sense of age. What leaves first is the unflattering geometry: uneven eye spacing, the specific width of the nose, a mouth that sits slightly off-center, the texture of skin that has lived. Those details make a face belong to one person, and they are what a model has been rewarded for sanding away.
You will notice it in the eyes before anything else. The color might be close, but the gaze is generic, the lower lid too taut, the inner corner too neat. Then the mouth: fuller, more symmetrical, sitting in a face pulled toward whatever the training data treated as desirable. Mood lighting can hide this, which is why a drifted result can still feel like a success until you set it next to a real photograph.
Drift compounds. If you take a slightly idealized output and feed it back as the next reference, the next generation starts from an already softened identity. You end up with a character who shares a haircut with you and nothing else. The model is doing what it was asked: make this look better. If you would not recognize this person from the side, in bad light, without the styling, the identity did not survive.
Why enhance and beautiful cause the swap
The words that feel like quality instructions are often identity instructions in disguise. "Enhance" does not mean sharpen the existing face. In a portrait model it usually means move the face toward the model's idea of a good face. "Beautiful," "stunning," "gorgeous," "perfect skin," "model-like," and "glamorous" do the same work. They tell the model that attractiveness is the objective and that the current face is allowed to change in service of it.
Those words sit on the same axis as the training signal. Images captioned with beauty language are a filtered sample of human faces, not a representative one. When you add that language to a prompt, you are not asking the model to light your face well. You are asking it to retrieve from the beauty cluster. Your reference image becomes a starting hint, not a lock.
Milder words leak too. "Improve," "fix," "clean up," "make it pop," and "more professional" can all be read as permission to regularize the face. Pores, faint lines, and asymmetry get treated as defects, and the model produces a face that would pass in a catalog and fail a family member.
A prompt is a pile of weighted requests. Beauty adjectives are high-weight, well-trained requests. Identity, unless you spell it out as a hard constraint, is a weaker request inferred from pixels. The model resolves the conflict by keeping the pose and replacing the person. If you want a better photograph of the same person, name the photographic change: harder light, a different lens feel, a cooler grade, a new location. Do not ask the face to become more itself by becoming more beautiful.
Restyling a room is not restyling a face
Background restyle and person restyle are different jobs, and mixing them in one prompt is how likeness dies.
A background can change freely: the wall, the weather, the city, the color of the light bouncing into the scene. The model can invent a lot there without touching identity, provided you tell it that the person is not the site of invention.
A person restyle is a change to hair, makeup, expression, age, body, or facial structure. A new jacket is usually compatible with likeness. A new hairstyle can be, if the face is locked. A change to bone structure never is, even if you call it subtle. Makeup preserves identity when you describe it as paint on an unchanged face, and destroys it when you describe it as a glow-up.
The failure mode is a prompt that sounds like a scene change and is actually a character change. "Cinematic portrait in golden hour" will warm the light and, unless blocked, the face. Write the prompt as two zones. Zone one is the person: who they are, what must not move. Zone two is the world: what is allowed to move. If a word could apply to either, put it in the world zone. "Cinematic" should modify lighting and grade, not the face. "Fashion" should modify clothes, not bone structure.
If you do want to change the person — a haircut, a costume, a period look — say so in a separate clause and still lock the facial geometry. Costume is clothing. Period is styling. Neither is a license to rebuild the skull.
Preservation clauses that hold a likeness
A preservation clause tells the model what is not a creative opportunity. "Keep the same person" is better than nothing and still weak, because the model can keep a same-ish person. You need to name the parts of identity that beauty language usually erases.
Lock structure before surface. Say that facial geometry stays: the distance between the eyes, the width and shape of the nose, the mouth, the jaw, the cheekbones, the brow. Say that unique marks stay: moles, scars, asymmetry, the real hairline, the real ear shape. Say that skin is not to be smoothed into a different person, and that age stays.
Be explicit about what "same" means: same identity, same facial structure, same proportions, same distinguishing features, no beautification, no symmetry correction, no skin smoothing that removes character, no change to eye shape or size. If you do not forbid idealization, idealization is available.
Order is part of the mechanism. Models attend strongly to early instructions and to repeated constraints. A working structure looks like this. First, the identity lock. Second, any allowed change to the person, limited to hair, clothes, or expression. Third, the scene and camera. Fourth, a short restatement of the lock. A prompt that spends most of its length on atmosphere will spend most of its attention there. Repeating the lock after the scene is a counterweight, not decoration.
Do not combine a preservation clause with a beauty clause and hope they cancel. "Keep her exact face, make her more beautiful" is a contradiction. The model will pick the easier, more rewarded half. If you need flattering light, ask for flattering light. Light can be kind without rebuilding the nose. The reference image is part of the clause, not a substitute for it. The image shows the face. The clause tells the model that the image is a rule rather than a mood board.
When a reference image beats any amount of prose
You cannot describe a face well enough. You can describe it at length and still miss the small spatial fact that makes it yours. Prose is good at style, setting, mood, and camera. Prose is bad at identity because identity is a spatial pattern, not a list of adjectives. "Brown eyes, dark hair, oval face" matches a crowd. The reference matches one.
Use a reference whenever likeness is the point. Prefer a clean, well-lit, front-facing or three-quarter image, because the model can only preserve what it can see. A tiny, occluded, heavily filtered source gives the model license to invent the missing parts, and invention is where drift lives. Sunglasses, beauty filters, and extreme angles hide the geometry you are trying to keep.
One strong reference is worth more than a paragraph of facial description. Extra prose should specify what may change around the person, not reconstruct the person in words. Tell the model to treat the image as identity, not inspiration.
In PicX Studio, as in any image-to-image portrait flow, the reference is the identity document and the prompt is the brief for everything else. Treat them that way and the model has a chance. Treat the prompt as a complete description of a person and the reference as a loose suggestion, and you will get a competent stranger.
How to check the face you actually got
The last failure is not the model. It is you deciding the result is close enough because the picture is attractive.
Do not judge likeness from the generated image alone. Put the source and the result side by side. Cover the hair and the clothes. Look at the middle of the face: the gap between the eyes, the root of the nose, the philtrum, the corners of the mouth, the thickness of the upper lid. If those moved, identity moved. A matching hairstyle is not a matching person.
Ask someone who knows the face, without showing them the source first. If they cannot name the subject, the picture failed, however good it looks. You are a biased judge because you know what you meant. You will fill in identity from memory.
Crop to just the eyes, then just the mouth. Drift hides in a well-composed medium shot and becomes obvious in a crop. Watch for the telltale upgrades: larger irises, higher cheekbones, a narrower nose, a more even hairline, skin without pores, a jaw that looks engineered. Those are not photographic improvements. They are substitutions.
If you asked for a new background and received a new skeleton, start over with a tighter preservation block and fewer beauty words. Do not iterate on the drifted frame. Iterating on a stranger produces a more polished stranger. Keep the original reference in the comparison every round, because your memory of a face adapts toward the last pretty version you saw. If the scene is right and the face is wrong, reuse the world's half of the prompt and rewrite the person block. If the face is right and the scene is wrong, change only the world. Mixing both again is how you lose the one thing you already had.



